High Level Design

Designing a Load Balancer

End-to-end design of a production-grade load balancer — control and data planes, health checking, session persistence, SSL termination, and high-availability with active-active failover.

August 10, 2026

A Load Balancer distributes incoming network traffic across multiple servers so no single server becomes overwhelmed. Instead of routing all traffic to one server (which eventually crashes under heavy load), a load balancer acts as a traffic cop — intelligently spreading requests across a pool of healthy servers.


Requirements#

Functional#

  • Traffic Distribution — distribute requests across backend servers using configurable algorithms.
  • Health Checking — continuously monitor backends and auto-remove unhealthy ones from the pool.
  • Session Persistence — support sticky sessions to route requests from the same client to the same server.
  • SSL Termination — handle TLS encryption/decryption to offload work from backend servers.
  • Layer 4 and Layer 7 — support both transport-level (TCP/UDP) and application-level (HTTP/HTTPS) load balancing.

Non-Functional#

  • High Availability — 99.99% uptime, no single point of failure.
  • Low Latency — < 1ms added overhead per request.
  • High Throughput — up to 1 million requests per second at peak.
  • Scalability — horizontal scaling to handle increasing traffic.
  • Fault Tolerance — continue operating when individual components fail.

Estimations#

CategoryMetricValue
Peak RPSMax load1,000,000
Average RPS~⅓ of peak~300,000
Concurrent connectionsRPS × avg duration (0.5s)~500,000
Ingress bandwidth1M × 2 KB2 GB/s
Egress bandwidth1M × 10 KB10 GB/s
Total bandwidth—~12 GB/s (~96 Gbps)
Memory per connection~500 bytes × 500K connections~250 MB
Health checks1000 servers / 5s interval200 checks/sec

Primary bottleneck: network bandwidth (~96 Gbps) — requires horizontal scaling.


Core APIs (Control Plane)#

The system splits into two planes:

PlanePurposeCharacteristics
Data PlaneHandles actual request trafficHigh throughput, low latency
Control PlaneConfiguration & managementLow traffic, explicit APIs

plane architecture

1. Register Backend Server#

POST /backends

Adds a server to the pool and starts health checking it. Initial status is "unknown" — transitions to "healthy" or "unhealthy" after the first health check.

register backend API

2. Remove Backend Server#

DELETE /backends/{backend_id}

Decommissions a server — stops new traffic immediately, allows existing connections to complete gracefully (connection draining).

remove backend API

3. Get Backend Health Status#

GET /backends/{backend_id}/health

Returns detailed health info including current state and historical metrics — useful for debugging and monitoring dashboards.

health check API

4. Configure Load Balancing Algorithm#

PUT /config/algorithm

Changes traffic distribution algorithm. Takes effect immediately for new connections; existing connections are not affected.

configure algorithm API


High Level Design#

The architecture splits into two parts:

  • Data plane — handles actual traffic at high speed; every request flows through it.
  • Control plane — manages configuration and health checking (slower pace; changes are infrequent).

architecture

1. Traffic Distribution#

Three components handle request routing:

data plane components

Frontend Listeners — bind to ports (80/443), accept TCP connections, parse enough of the request for routing decisions (headers, URL path for L7), and hand off to the routing engine.

Routing Engine — the brain of the LB. Maintains the backend list, health state, connection counters, and configured algorithm. Selects the target backend for every incoming connection.

Backend Pool — a logical group of interchangeable servers.

  • Simple: single pool, all requests routed to one group.
  • Complex: multiple pools by service type — API pool, static content pool, auth pool.

Request Flow#

request flow

  1. Client sends a request to the LB's public IP.
  2. Frontend listener accepts the TCP connection; parses request metadata for L7.
  3. Routing engine filters unhealthy backends, applies algorithm, selects target.
  4. Request forwarded to selected backend.
  5. Backend processes and responds; LB forwards response to client.
  6. Connection kept alive (HTTP keep-alive) or closed per configuration.

2. Health Monitoring#

Without health checking, the LB blindly sends traffic to dead servers — users see errors. The Health Checker is a background process that:

  • Probes each backend every 5–10 seconds.
  • Tracks successive successes and failures.
  • Marks unhealthy after 2–3 consecutive failures; re-enables after 2–3 consecutive successes.
  • Notifies the routing engine on every status change.
TypeHow it worksWhen to use
TCP CheckOpens + closes TCP connectionBasic connectivity
HTTP CheckGET /health → expects 2xxWeb apps / APIs
Custom ScriptUser-defined scriptComplex validation, legacy systems

Most production deployments use HTTP health checks — the application exposes /health or /healthz that returns 200 OK only when it can actually handle requests.


3. High Availability#

Problem: the LB is now a single point of failure.

Solution: run multiple LB instances.

AspectActive-PassiveActive-Active
Traffic handlingOnly one node serves trafficAll nodes serve simultaneously
Client connectionVia VIP (Virtual IP)Multiple IPs or shared IP
FailoverStandby takes over VIP (1–3s)Other nodes continue, no downtime
Resource utilization❌ 50% wasted (idle standby)✅ Fully utilized
State managementNo sync neededRequires shared state store
ScalabilityLimitedHighly scalable

Active-active is the better choice for production. Full resource utilization and instant failover outweigh the added complexity of state sync.

high level design

How it works:

  1. Clients connect to the VIP — they never know which LB node handles them.
  2. Both LB nodes serve traffic simultaneously (active-active).
  3. Each LB node routes only to healthy backends (unhealthy ones are excluded).
  4. Sticky session state stored in Redis — any LB node can look up which backend a user belongs to.
  5. Health Checker monitors all backends and both LB nodes continuously.
  6. Config Manager pushes updates to all LB nodes from a central Config Store (etcd/Consul).

Component Summary#

ComponentRole
Virtual IP / DNSStable entry point; abstracts the LB cluster
LB NodesAccept requests; route to backends — stateless, horizontally scalable
Session Store (Redis)Shared sticky session mappings across LB nodes
Health CheckerBackground process; ensures only healthy nodes receive traffic
Config ManagerManages LB config via APIs; handles updates and validation
Config StorePersists config (etcd, Consul, or PostgreSQL)
Backend PoolApplication servers; organized by service type

Database Design#

Data TypeStorageWhy
Active connectionsIn-memory (per LB node)Microsecond latency; local ownership
Backend health statusIn-memory (per LB node)Frequently updated; fast routing access
Connection countersIn-memory (per LB node)Updated every request (least-connections)
Session mappingsRedis clusterShared across LB nodes for sticky sessions
Backend configurationetcd / Consul / PostgreSQLPersistent, versioned, survives restarts
Metrics & logsPrometheus / InfluxDBHistorical analysis and alerting

Hot path must never touch disk or cross a network boundary for core routing. Config is loaded into memory at startup and updated via push notifications.


Deep Dive#

Load Balancing Algorithms#

AlgorithmComplexityStateAdaptiveBest For
Round RobinVery simpleStatelessNoneHomogeneous backends, uniform requests
Weighted Round RobinSimpleStatelessNoneKnown capacity differences, gradual rollouts
Least ConnectionsModeratePer-node stateHighVariable request processing times
Weighted Least ConnectionsModeratePer-node stateHighProduction systems (default recommendation)
IP HashSimpleStatelessNoneBasic session persistence, non-HTTP protocols
Consistent HashingComplexStatelessNoneDynamic scaling, cache servers

Recommendations:

  • Default to Weighted Least Connections — handles heterogeneous backends and variable request costs.
  • For session persistence, prefer cookie-based stickiness over IP hash (IP hash is fragile due to NAT).
  • Use Consistent Hashing when backends scale frequently or when load balancing cache servers.

Session Persistence (Sticky Sessions)#

Problem: stateful applications store session data in a specific backend's memory. A subsequent request routed to a different backend looks like a logout.

The LB injects a cookie identifying the target backend. Every subsequent request includes the cookie.

cookie-based stickiness

✅ Works regardless of client IP changes (mobile, corporate NAT)
✅ Survives LB restarts
✅ No shared state between LB nodes

Approach 2: Source IP Persistence#

Routes based on hash(client_ip) % n. Use only for non-HTTP protocols or when cookies are not possible. Breaks for mobile users and corporate NAT.

Approach 3: Externalized Session Store (Best Long-Term)#

Store session data in a shared Redis cluster — makes backends fully stateless.

centralized session store

✅ No sticky sessions needed — any LB algorithm works
✅ Backends scale freely
✅ Backend failures don't cause session loss

Always build new applications stateless. Use externalized session store.


Layer 4 vs Layer 7#

AspectLayer 4 (TCP/UDP)Layer 7 (HTTP/HTTPS)
Routing basisIP + PortURL, headers, cookies, body
Performance⚡ Very high (packet-level)Moderate (request inspection)
Content awareness❌✅
SSL termination❌✅
Caching / Auth / Rate Limiting❌✅
ExamplesAWS NLB, HAProxyNGINX, AWS ALB

SSL/TLS Termination#

LB handles all encryption; backends receive plain HTTP.

Benefits:

  • One place to install/renew certificates (not on every backend).
  • Offloads CPU-intensive cryptography from backends.
  • Enables content-based routing (URL/header inspection requires decrypted traffic).
  • TLS session caching — repeat visitors skip the full handshake.

Security trade-off: backend ↔ LB connection is unencrypted.

  • Inside a private VPC → acceptable.
  • For end-to-end encryption → re-encryption (LB decrypts, inspects, re-encrypts) or SSL passthrough (LB routes based on SNI, no content inspection).

Failure Handling#

Detection mechanisms:

  • Heartbeat messages — LB nodes exchange "I'm alive" every 1–3 seconds.
  • Health checks — same mechanism used for backends, applied to LB nodes.
  • Shared storage heartbeats — write timestamps to etcd/Redis to detect liveness.

Balance sensitivity: too sensitive → false positives (brief network hiccup declared dead); too slow → extended outage.

Graceful degradation (Connection Draining):

  1. Detection — health check fails.
  2. Draining — stop new traffic; allow existing connections to complete.
  3. Removal — fully remove backend once in-flight requests finish.
  4. Recovery — add backend back once it passes the success threshold.