High Level Design

Kafka

Kafka's architecture, core components, partitions, ISR-based durability, trade-offs, and when to use it over RabbitMQ or SQS.

August 10, 2026

A distributed, durable, and scalable event-streaming platform used to publish, store, and consume streams of records in real time.

Pain Points Kafka Solves#

  • Synchronous APIs cause high latency and cascading failures
  • Direct service-to-service communication leads to tight coupling
  • Traditional message queues struggle with scale, durability, and replay
  • Systems need event replay for debugging, analytics, and recovery

Use Cases#

  • Event-driven microservices
  • Activity logs, audit logs
  • Data ingestion pipelines
  • Replayable event streams

Architecture#

Kafka architecture

Core Components#

ComponentRole
ProducerPublishes messages
TopicLogical stream of messages — durable
PartitionOrdered, append-only log within a topic
BrokerKafka server storing partitions
ConsumerReads messages
Consumer GroupEnables parallel consumption
OffsetConsumer's read position within a partition
ZooKeeper / KRaftMetadata and cluster coordination

Flow#

  1. Producer publishes events to a Topic
  2. Topic is split into Partitions (ordered logs)
  3. Partitions are distributed across Brokers
  4. Events are persisted on disk and replicated
  5. Consumers read events using offsets
  6. Events are retained for a configured retention period — not consumed then deleted

Key idea: Kafka is a pull-based, log-centric system — not a traditional queue.


Trade-offs#

ProsCons
Extremely high throughputOperational complexity
Horizontal scalabilityEventual consistency by default
Durable storage with replayMessage ordering only within a partition
Decoupled architectureNot ideal for low-latency request/response
Fault tolerant via replication

Scalability#

  • Write scaling → Add partitions
  • Read scaling → Add consumers in a group
  • Broker scaling → Add brokers horizontally
  • Near-linear throughput scaling with partitions

Caution: Too many partitions → metadata + consumer rebalance overhead.


Consistency & Availability#

  • Ordering: Guaranteed per partition only
  • Availability: Replication + leader election + ISR
  • CAP view: Prioritizes Availability + Partition tolerance; consistency is configurable

ISR — In-Sync Replicas#

The set of replicas fully caught up with the leader.

ISR diagram

Durability via acks + ISR#

SettingBehavior
acks=allProducer waits until all ISR replicas confirm. Strong durability.
acks=1Only leader writes. Lower durability.

min.insync.replicas#

Defines how many ISR replicas must acknowledge a write:

  • Replication factor = 3, min.insync.replicas = 2
  • Guarantee: at least 2 replicas persist the data
  • If ISR shrinks below 2 → writes are rejected (prevents data loss on leader failure)

ISR is the key durability safety valve.


Kafka vs Alternatives#

ComparisonKafkaAlternative
vs RabbitMQHigh throughput, replay, streamingLow latency, routing, queues
vs SQSSelf-managed, high control, replayFully managed, limited replay

Real-World Usage#

  • Netflix — event pipelines and monitoring
  • Uber — real-time analytics
  • LinkedIn — activity streams (Kafka originated here)
  • Fintech — transaction events and audit logs