High Level Design
Handling Numbers in System Design
The essential reference for system design interviews — powers of two, latency numbers every engineer should know, and availability nines decoded.
Every system design interview involves estimation. Before you can reason about scale, you need a mental model of the numbers: how big is a gigabyte, how fast is a disk seek, how much downtime does "four nines" allow? This post is that mental model.
Power of Two#
Although data volumes in distributed systems can become enormous, the math always comes back to the basics.
KMGTP → Kids Make Good Team Players A mnemonic for: Kilo, Mega, Giga, Tera, Peta
A byte is a sequence of 8 bits. An ASCII character uses one byte of memory. All storage and memory sizes are expressed as powers of 2 — because computers use binary (0 or 1), and memory hardware is designed in doublings.
| Power | Approximate Value | Full Name | Short Name |
|---|---|---|---|
| 2^10 = 1,024 ≈ 10^3 | 1 Thousand | 1 Kilobyte | 1 KB |
| 2^20 ≈ 10^6 | 1 Million | 1 Megabyte | 1 MB |
| 2^30 ≈ 10^9 | 1 Billion | 1 Gigabyte | 1 GB |
| 2^40 ≈ 10^12 | 1 Trillion | 1 Terabyte | 1 TB |
| 2^50 ≈ 10^15 | 1 Quadrillion | 1 Petabyte | 1 PB |
Why powers of 2? Computers are built on binary logic — every bit is 0 or 1. Memory chips double in capacity each generation, so all sizes (RAM sticks: 4 GB, 8 GB, 16 GB; SSDs: 256 GB, 512 GB, 1 TB) are naturally powers of 2.
Quick size intuitions#
| Object | Size |
|---|---|
| ASCII character | 1 byte |
| UTF-8 character (most languages) | 1–4 bytes |
| Integer (32-bit) | 4 bytes |
| Long / Double (64-bit) | 8 bytes |
| UUID / GUID | 16 bytes |
| A tweet (280 chars) | ~280 bytes |
| A typical web page (HTML) | ~50–200 KB |
| A high-res photo (JPEG) | ~3–5 MB |
| A 1-hour video (compressed) | ~700 MB – 2 GB |
| A full-length movie (4K) | ~50–100 GB |
Latency Numbers Every Engineer Should Know#
These numbers come from Jeff Dean's famous 2010 research (updated for modern hardware). The exact values change year to year, but the relative magnitudes stay constant — and that's what matters for design conversations.
| Operation | Latency | Notes |
|---|---|---|
| L1 cache reference | 0.5 ns | |
| Branch mispredict | 5 ns | |
| L2 cache reference | 7 ns | 14× slower than L1 |
| Mutex lock/unlock | 100 ns | |
| Main memory reference | 100 ns | 20× slower than L2 |
| Compress 1 KB (Snappy) | 10 µs | |
| Send 1 KB over 1 Gbps network | 10 µs | |
| Read 4 KB randomly from SSD | 150 µs | |
| Read 1 MB sequentially from memory | 250 µs | |
| Round trip within same datacenter | 500 µs | |
| Read 1 MB sequentially from SSD | 1 ms | 4× slower than memory |
| Disk seek (HDD) | 10 ms | 20× slower than SSD |
| Read 1 MB sequentially from HDD | 20 ms | |
| Send packet: California → Netherlands → California | 150 ms |
Conclusions to draw in an interview#
| Takeaway | What it means for design |
|---|---|
| Memory is fast, disk is slow | Cache aggressively; avoid hitting disk on the hot path |
| Avoid disk seeks | Sequential I/O is 20× faster than random; prefer append-only logs |
| Simple compression is worth it | 10 µs to compress saves more time than sending uncompressed over the network |
| In-datacenter RTT ≈ 500 µs | Services in the same region can talk cheaply; cross-region adds 150ms+ |
| Memory is ~10,000× faster than disk | Justify Redis / memcached with this number |
Availability Numbers#
Availability is the percentage of time a system is operational and accessible. It's often expressed as "nines."
Availability = Uptime / (Uptime + Downtime)
The nines table#
| Availability | Nines | Downtime / Year | Downtime / Month | Downtime / Week |
|---|---|---|---|---|
| 90% | 1 nine | 36.5 days | 72 hours | 16.8 hours |
| 99% | 2 nines | 3.65 days | 7.2 hours | 1.68 hours |
| 99.9% | 3 nines | 8.77 hours | 43.8 minutes | 10.1 minutes |
| 99.99% | 4 nines | 52.6 minutes | 4.38 minutes | 1.01 minutes |
| 99.999% | 5 nines | 5.26 minutes | 26.3 seconds | 6.05 seconds |
| 99.9999% | 6 nines | 31.5 seconds | 2.63 seconds | ~1 second |
Interview tip: When an interviewer asks for "high availability," they usually mean 99.9% (3 nines) or better. Pushing to 5 nines requires significant complexity — redundant components, active-active failover, zero-downtime deploys. Always ask what the SLA target is; it drives the whole architecture.
Why availability compounds#
When you chain services together, overall availability multiplies:
| Setup | Availability |
|---|---|
| Two services, both 99.9% | 99.9% × 99.9% = 99.8% |
| Three services, all 99.9% | 99.9%³ = 99.7% |
| Five services, all 99.9% | 99.9%⁵ = 99.5% |
This is why microservices need careful reliability budgets — more hops = more ways to fail.
To improve availability:
- Redundancy — no single point of failure; replicate every critical component
- Failover — automatic promotion of a standby (active-passive or active-active)
- Health checks + circuit breakers — stop sending traffic to a failing node before humans notice
- Graceful degradation — return partial results rather than a hard error
Back-of-Envelope Estimation Cheat Sheet#
Traffic#
| DAU | QPS (reads) | Peak QPS (×3) |
|---|---|---|
| 1 million | ~12 | ~36 |
| 10 million | ~115 | ~350 |
| 100 million | ~1,150 | ~3,500 |
| 1 billion | ~11,500 | ~35,000 |
Formula: QPS = DAU × (requests/user/day) ÷ 86,400 seconds
Storage#
| Scenario | Quick math |
|---|---|
| 10M users, each stores 1 MB/day | 10 TB/day |
| 500K video uploads/day, 300 MB avg | 150 TB/day |
| 100M tweets/day, 280 bytes each | ~28 GB/day |
| 1B photos/day, 300 KB each | 300 TB/day |
Rule of thumb: At 1M users storing 1 KB/day → 1 GB/day. Scale linearly from there.
Network bandwidth#
| Scenario | Bandwidth |
|---|---|
| 1 Gbps link | 125 MB/s |
| Serving 100 MB/s of video | 800 Mbps |
| 10,000 QPS, 1 KB response each | ~10 MB/s = 80 Mbps |
Useful constants#
| Value | Number |
|---|---|
| Seconds in a day | 86,400 (~10^5) |
| Seconds in a month | ~2.6 million (~2.5 × 10^6) |
| Seconds in a year | ~31.5 million (~3 × 10^7) |
| 1 Gbps = | 125 MB/s |
Common Estimation Mistakes to Avoid#
| Mistake | Fix |
|---|---|
| Forgetting peak vs average traffic | Assume peak is 2–3× average |
| Ignoring replication overhead | Storage × replication factor (usually 3× for distributed systems) |
| Confusing MB and MB/s | State units explicitly every time |
| Treating 99% and 99.9% as "basically the same" | They differ by 8 hours of downtime per year |
| Ignoring write amplification | Writes to replicated systems multiply: 1 write → N disk writes |