High Level Design

Data Replication & Migration

Primary-replica replication, WAL and CDC, async vs sync vs semi-sync trade-offs, split-brain problem, write amplification, and zero-downtime migration strategies.

August 10, 2026

Replication is the process of creating and maintaining multiple copies of the same data on different servers (replicas).


Why Replication?#

GoalDescriptionExample
High AvailabilityKeep service running when one node failsPrimary DB crashes → replica takes over
Read ScalabilityDistribute read loadMultiple app servers query replicas
Disaster RecoveryCross-region backupsUS → EU replication
Fault ToleranceSurvive hardware or network failuresLeader crash → follower promotion
Geo LatencyServe users from nearby replicasIndia users read from APAC replica

Primary-Replica Architecture#

  • Writes → only to the primary (maintains consistency)
  • Reads → from both primary and replicas
  • Especially beneficial for read-intensive applications

WAL — Write-Ahead Log#

Replicas process write operations sequentially from the primary's log:

  • ✅ Transfers only necessary operations (efficient)
  • ❌ If replica can't keep up → consistency issues
  • ❌ Timestamps and contextual values can cause confusion

CDC — Change Data Capture#

Sends events indicating data changes; subscribers process these events and make necessary transformations:

  • Useful when you have different DB types (write-optimized primary, read-optimized replica)
  • Built-in libraries for connecting to various databases
  • Simplifies replication across heterogeneous systems

Types of Replication#

TypeWrite AcknowledgementProsConsUse Cases
AsynchronousPrimary does NOT wait for replicasLow latency, high throughputReplication lag, data loss riskSocial feeds, analytics
SynchronousPrimary waits for ALL replicasStrong consistency, no data lossHigh latency, lower availabilityBanking, finance
Semi-SynchronousPrimary waits for at least ONE replicaBalanced consistency & latencySome lag, complex setupE-commerce, user profiles

Quick Comparison#

TypeLatencyConsistencyData Loss RiskAvailability
AsyncLowWeakHighHigh
SyncHighStrongNoneLower
Semi-syncMediumMediumLowMedium

Interview tip: Choose Async for performance, Sync for correctness, Semi-sync for balanced systems.


Challenges#

Split-Brain Problem#

Two nodes both believe they are the primary, leading to write conflicts and inconsistencies.

  • Solution: Use an odd number of primary nodes to ensure a majority consensus
  • Even number of nodes → tie votes → manual reconciliation required

Write Amplification#

More data is written to storage than the original amount intended. Puts heavy load on primary DB bandwidth.

Solution: Use a consensus mechanism like Paxos or Raft to ensure the entire cluster agrees on a value.

Write amplification diagram


Database Migration#

Reasons to migrate: infrastructure upgrade, cloud adoption, system modernization.

Naive Approach (Downtime Required)#

  1. Stop the existing DB (all incoming requests fail)
  2. Take a full data dump into the new DB
  3. Point servers to the new DB (requires deployment/restart)

Problem: Significant downtime

Optimised Approach (Zero-Downtime)#

  1. Set up Change Data Capture (CDC) on the old DB — tracks all INSERT/UPDATE/DELETE
  2. Take a full snapshot of existing data
  3. Set up a proxy in front of both DBs to serve existing clients
  4. Point clients to the proxy
  5. Point the proxy to the new DB
  6. Point clients directly to the new DB (final deployment)
  7. Remove the proxy

Challenge: Updates during migration require either temporarily blocking updates or ensuring all clients switch simultaneously to avoid inconsistency.