High Level Design

System Design Trade-offs

Pull vs Push, Memory vs Latency, Throughput vs Latency, Consistency vs Availability, Latency vs Accuracy — every key trade-off with decision frameworks.

August 10, 2026

Understanding trade-offs is the core skill of system design. Every architectural decision involves giving something up to gain something else.


Pull vs Push Architecture#

AspectPullPush
Who initiatesClient initiates the data requestServer pushes data to client
LatencyHigher (depends on user action)Lower (real-time updates)
ComplexitySimpler to implement and scaleMore complex (managing connections/subscriptions)
Resource UsageEfficient server usageHeavier server/network load
Failure HandlingEasier to handle offline clientsHarder — need retries, reconnections
ExamplePull-to-refresh feedPush notifications, real-time messaging

Hybrid Model (Best of Both)#

Use a message queue to decouple:

AspectHybrid (Message Queue)
Core IdeaServer pushes events to queue; clients pull from or are notified via queue
ExampleDM service → Kafka topic → Notification service → Push notification
ScalabilityHighly scalable; decouples producers and consumers
ReliabilityRetries, DLQ, acknowledgements

Why use hybrid? Real-time push for alerts + on-demand pull for feeds + loose coupling + fault tolerance.


Memory vs Latency#

AspectHigh Memory (Caching)Low Memory
LatencyVery low (~10ms for cached data)Higher (DB or remote fetch)
Resource UsageMore RAM, higher costLess memory, cheaper
ConsistencyMay serve stale dataAlways latest (from source)
Failure ResilienceCache-aside works if DB goes downComplete failure if backend down

Balanced approach: Keep hot data in cache with LRU/LFU eviction + tiered storage (Redis → DB fallback).


Throughput vs Latency#

AspectHigh ThroughputLow Latency
FocusMaximize system capacityMinimize time per request
Suitable forBatch processing, analyticsReal-time UX, interactive apps
OptimizationQueues, batch jobs, parallelismIn-memory cache, fast DBs, fast networks
Trade-offAdding throughput can add queuing delayLow latency may limit concurrent capacity

Balanced approach: Asynchronous processing — fast response for the user, high throughput in the background.


Consistency vs Availability (CAP)#

AspectConsistencyAvailability
DefinitionAll nodes return the most recent writeSystem serves requests even if some nodes fail
User ExperienceAccurate, up-to-date dataResponsive, even if data might be slightly outdated
Failure ScenarioMay reject requests if nodes don't agreeMay serve stale data but won't reject
Suitable forBanking, inventory, ordersSocial media, content delivery

Balanced approach: Use eventual consistency where strong consistency isn't critical. Design fallback mechanisms (show cached feed if DB is unreachable).


Latency vs Accuracy#

AspectLatencyAccuracy
DefinitionTime delay between input and outputHow correct/precise the result is
Measured inMillisecondsPercentage (%, precision/recall)
Critical inReal-time systems, gaming, HFTMedical diagnosis, fraud detection
OptimizationModel compression, edge compute, cachingData quality, larger models, advanced training

Reducing latency often involves making educated guesses from cached/past data — trading exactness for speed.


Trade-off Summary Matrix#

DimensionIncreasesDecreases
MemoryLowers latency ↓Raises cost ↑
Throughput—Latency ↑ (inversely related)
CostMay reduce latency—
LatencyTolerating inconsistency allows faster response—
Consistency—Availability ↓ (CAP theorem)
Availability—Consistency ↓