High Level Design
Kafka
Kafka's architecture, core components, partitions, ISR-based durability, trade-offs, and when to use it over RabbitMQ or SQS.
A distributed, durable, and scalable event-streaming platform used to publish, store, and consume streams of records in real time.
Pain Points Kafka Solves#
- Synchronous APIs cause high latency and cascading failures
- Direct service-to-service communication leads to tight coupling
- Traditional message queues struggle with scale, durability, and replay
- Systems need event replay for debugging, analytics, and recovery
Use Cases#
- Event-driven microservices
- Activity logs, audit logs
- Data ingestion pipelines
- Replayable event streams
Architecture#

Core Components#
| Component | Role |
|---|---|
| Producer | Publishes messages |
| Topic | Logical stream of messages — durable |
| Partition | Ordered, append-only log within a topic |
| Broker | Kafka server storing partitions |
| Consumer | Reads messages |
| Consumer Group | Enables parallel consumption |
| Offset | Consumer's read position within a partition |
| ZooKeeper / KRaft | Metadata and cluster coordination |
Flow#
- Producer publishes events to a Topic
- Topic is split into Partitions (ordered logs)
- Partitions are distributed across Brokers
- Events are persisted on disk and replicated
- Consumers read events using offsets
- Events are retained for a configured retention period — not consumed then deleted
Key idea: Kafka is a pull-based, log-centric system — not a traditional queue.
Trade-offs#
| Pros | Cons |
|---|---|
| Extremely high throughput | Operational complexity |
| Horizontal scalability | Eventual consistency by default |
| Durable storage with replay | Message ordering only within a partition |
| Decoupled architecture | Not ideal for low-latency request/response |
| Fault tolerant via replication |
Scalability#
- Write scaling → Add partitions
- Read scaling → Add consumers in a group
- Broker scaling → Add brokers horizontally
- Near-linear throughput scaling with partitions
Caution: Too many partitions → metadata + consumer rebalance overhead.
Consistency & Availability#
- Ordering: Guaranteed per partition only
- Availability: Replication + leader election + ISR
- CAP view: Prioritizes Availability + Partition tolerance; consistency is configurable
ISR — In-Sync Replicas#
The set of replicas fully caught up with the leader.

Durability via acks + ISR#
| Setting | Behavior |
|---|---|
| acks=all | Producer waits until all ISR replicas confirm. Strong durability. |
| acks=1 | Only leader writes. Lower durability. |
min.insync.replicas#
Defines how many ISR replicas must acknowledge a write:
- Replication factor = 3, min.insync.replicas = 2
- Guarantee: at least 2 replicas persist the data
- If ISR shrinks below 2 → writes are rejected (prevents data loss on leader failure)
ISR is the key durability safety valve.
Kafka vs Alternatives#
| Comparison | Kafka | Alternative |
|---|---|---|
| vs RabbitMQ | High throughput, replay, streaming | Low latency, routing, queues |
| vs SQS | Self-managed, high control, replay | Fully managed, limited replay |
Real-World Usage#
- Netflix — event pipelines and monitoring
- Uber — real-time analytics
- LinkedIn — activity streams (Kafka originated here)
- Fintech — transaction events and audit logs