High Level Design
System Design Trade-offs
Pull vs Push, Memory vs Latency, Throughput vs Latency, Consistency vs Availability, Latency vs Accuracy — every key trade-off with decision frameworks.
Understanding trade-offs is the core skill of system design. Every architectural decision involves giving something up to gain something else.
Pull vs Push Architecture#
| Aspect | Pull | Push |
|---|---|---|
| Who initiates | Client initiates the data request | Server pushes data to client |
| Latency | Higher (depends on user action) | Lower (real-time updates) |
| Complexity | Simpler to implement and scale | More complex (managing connections/subscriptions) |
| Resource Usage | Efficient server usage | Heavier server/network load |
| Failure Handling | Easier to handle offline clients | Harder — need retries, reconnections |
| Example | Pull-to-refresh feed | Push notifications, real-time messaging |
Hybrid Model (Best of Both)#
Use a message queue to decouple:
| Aspect | Hybrid (Message Queue) |
|---|---|
| Core Idea | Server pushes events to queue; clients pull from or are notified via queue |
| Example | DM service → Kafka topic → Notification service → Push notification |
| Scalability | Highly scalable; decouples producers and consumers |
| Reliability | Retries, DLQ, acknowledgements |
Why use hybrid? Real-time push for alerts + on-demand pull for feeds + loose coupling + fault tolerance.
Memory vs Latency#
| Aspect | High Memory (Caching) | Low Memory |
|---|---|---|
| Latency | Very low (~10ms for cached data) | Higher (DB or remote fetch) |
| Resource Usage | More RAM, higher cost | Less memory, cheaper |
| Consistency | May serve stale data | Always latest (from source) |
| Failure Resilience | Cache-aside works if DB goes down | Complete failure if backend down |
Balanced approach: Keep hot data in cache with LRU/LFU eviction + tiered storage (Redis → DB fallback).
Throughput vs Latency#
| Aspect | High Throughput | Low Latency |
|---|---|---|
| Focus | Maximize system capacity | Minimize time per request |
| Suitable for | Batch processing, analytics | Real-time UX, interactive apps |
| Optimization | Queues, batch jobs, parallelism | In-memory cache, fast DBs, fast networks |
| Trade-off | Adding throughput can add queuing delay | Low latency may limit concurrent capacity |
Balanced approach: Asynchronous processing — fast response for the user, high throughput in the background.
Consistency vs Availability (CAP)#
| Aspect | Consistency | Availability |
|---|---|---|
| Definition | All nodes return the most recent write | System serves requests even if some nodes fail |
| User Experience | Accurate, up-to-date data | Responsive, even if data might be slightly outdated |
| Failure Scenario | May reject requests if nodes don't agree | May serve stale data but won't reject |
| Suitable for | Banking, inventory, orders | Social media, content delivery |
Balanced approach: Use eventual consistency where strong consistency isn't critical. Design fallback mechanisms (show cached feed if DB is unreachable).
Latency vs Accuracy#
| Aspect | Latency | Accuracy |
|---|---|---|
| Definition | Time delay between input and output | How correct/precise the result is |
| Measured in | Milliseconds | Percentage (%, precision/recall) |
| Critical in | Real-time systems, gaming, HFT | Medical diagnosis, fraud detection |
| Optimization | Model compression, edge compute, caching | Data quality, larger models, advanced training |
Reducing latency often involves making educated guesses from cached/past data — trading exactness for speed.
Trade-off Summary Matrix#
| Dimension | Increases | Decreases |
|---|---|---|
| Memory | Lowers latency ↓ | Raises cost ↑ |
| Throughput | — | Latency ↑ (inversely related) |
| Cost | May reduce latency | — |
| Latency | Tolerating inconsistency allows faster response | — |
| Consistency | — | Availability ↓ (CAP theorem) |
| Availability | — | Consistency ↓ |