High Level Design
Data Replication & Migration
Primary-replica replication, WAL and CDC, async vs sync vs semi-sync trade-offs, split-brain problem, write amplification, and zero-downtime migration strategies.
Replication is the process of creating and maintaining multiple copies of the same data on different servers (replicas).
Why Replication?#
| Goal | Description | Example |
|---|---|---|
| High Availability | Keep service running when one node fails | Primary DB crashes → replica takes over |
| Read Scalability | Distribute read load | Multiple app servers query replicas |
| Disaster Recovery | Cross-region backups | US → EU replication |
| Fault Tolerance | Survive hardware or network failures | Leader crash → follower promotion |
| Geo Latency | Serve users from nearby replicas | India users read from APAC replica |
Primary-Replica Architecture#
- Writes → only to the primary (maintains consistency)
- Reads → from both primary and replicas
- Especially beneficial for read-intensive applications
WAL — Write-Ahead Log#
Replicas process write operations sequentially from the primary's log:
- ✅ Transfers only necessary operations (efficient)
- ❌ If replica can't keep up → consistency issues
- ❌ Timestamps and contextual values can cause confusion
CDC — Change Data Capture#
Sends events indicating data changes; subscribers process these events and make necessary transformations:
- Useful when you have different DB types (write-optimized primary, read-optimized replica)
- Built-in libraries for connecting to various databases
- Simplifies replication across heterogeneous systems
Types of Replication#
| Type | Write Acknowledgement | Pros | Cons | Use Cases |
|---|---|---|---|---|
| Asynchronous | Primary does NOT wait for replicas | Low latency, high throughput | Replication lag, data loss risk | Social feeds, analytics |
| Synchronous | Primary waits for ALL replicas | Strong consistency, no data loss | High latency, lower availability | Banking, finance |
| Semi-Synchronous | Primary waits for at least ONE replica | Balanced consistency & latency | Some lag, complex setup | E-commerce, user profiles |
Quick Comparison#
| Type | Latency | Consistency | Data Loss Risk | Availability |
|---|---|---|---|---|
| Async | Low | Weak | High | High |
| Sync | High | Strong | None | Lower |
| Semi-sync | Medium | Medium | Low | Medium |
Interview tip: Choose Async for performance, Sync for correctness, Semi-sync for balanced systems.
Challenges#
Split-Brain Problem#
Two nodes both believe they are the primary, leading to write conflicts and inconsistencies.
- Solution: Use an odd number of primary nodes to ensure a majority consensus
- Even number of nodes → tie votes → manual reconciliation required
Write Amplification#
More data is written to storage than the original amount intended. Puts heavy load on primary DB bandwidth.
Solution: Use a consensus mechanism like Paxos or Raft to ensure the entire cluster agrees on a value.

Database Migration#
Reasons to migrate: infrastructure upgrade, cloud adoption, system modernization.
Naive Approach (Downtime Required)#
- Stop the existing DB (all incoming requests fail)
- Take a full data dump into the new DB
- Point servers to the new DB (requires deployment/restart)
Problem: Significant downtime
Optimised Approach (Zero-Downtime)#
- Set up Change Data Capture (CDC) on the old DB — tracks all INSERT/UPDATE/DELETE
- Take a full snapshot of existing data
- Set up a proxy in front of both DBs to serve existing clients
- Point clients to the proxy
- Point the proxy to the new DB
- Point clients directly to the new DB (final deployment)
- Remove the proxy
Challenge: Updates during migration require either temporarily blocking updates or ensuring all clients switch simultaneously to avoid inconsistency.