High Level Design

Handling Numbers in System Design

The essential reference for system design interviews — powers of two, latency numbers every engineer should know, and availability nines decoded.

September 18, 2026·Updated September 23, 2026

Every system design interview involves estimation. Before you can reason about scale, you need a mental model of the numbers: how big is a gigabyte, how fast is a disk seek, how much downtime does "four nines" allow? This post is that mental model.


Power of Two#

Although data volumes in distributed systems can become enormous, the math always comes back to the basics.

KMGTP → Kids Make Good Team Players A mnemonic for: Kilo, Mega, Giga, Tera, Peta

A byte is a sequence of 8 bits. An ASCII character uses one byte of memory. All storage and memory sizes are expressed as powers of 2 — because computers use binary (0 or 1), and memory hardware is designed in doublings.

PowerApproximate ValueFull NameShort Name
2^10 = 1,024 ≈ 10^31 Thousand1 Kilobyte1 KB
2^20 ≈ 10^61 Million1 Megabyte1 MB
2^30 ≈ 10^91 Billion1 Gigabyte1 GB
2^40 ≈ 10^121 Trillion1 Terabyte1 TB
2^50 ≈ 10^151 Quadrillion1 Petabyte1 PB

Why powers of 2? Computers are built on binary logic — every bit is 0 or 1. Memory chips double in capacity each generation, so all sizes (RAM sticks: 4 GB, 8 GB, 16 GB; SSDs: 256 GB, 512 GB, 1 TB) are naturally powers of 2.

Quick size intuitions#

ObjectSize
ASCII character1 byte
UTF-8 character (most languages)1–4 bytes
Integer (32-bit)4 bytes
Long / Double (64-bit)8 bytes
UUID / GUID16 bytes
A tweet (280 chars)~280 bytes
A typical web page (HTML)~50–200 KB
A high-res photo (JPEG)~3–5 MB
A 1-hour video (compressed)~700 MB – 2 GB
A full-length movie (4K)~50–100 GB

Latency Numbers Every Engineer Should Know#

These numbers come from Jeff Dean's famous 2010 research (updated for modern hardware). The exact values change year to year, but the relative magnitudes stay constant — and that's what matters for design conversations.

OperationLatencyNotes
L1 cache reference0.5 ns
Branch mispredict5 ns
L2 cache reference7 ns14× slower than L1
Mutex lock/unlock100 ns
Main memory reference100 ns20× slower than L2
Compress 1 KB (Snappy)10 µs
Send 1 KB over 1 Gbps network10 µs
Read 4 KB randomly from SSD150 µs
Read 1 MB sequentially from memory250 µs
Round trip within same datacenter500 µs
Read 1 MB sequentially from SSD1 ms4× slower than memory
Disk seek (HDD)10 ms20× slower than SSD
Read 1 MB sequentially from HDD20 ms
Send packet: California → Netherlands → California150 ms

Conclusions to draw in an interview#

TakeawayWhat it means for design
Memory is fast, disk is slowCache aggressively; avoid hitting disk on the hot path
Avoid disk seeksSequential I/O is 20× faster than random; prefer append-only logs
Simple compression is worth it10 µs to compress saves more time than sending uncompressed over the network
In-datacenter RTT ≈ 500 µsServices in the same region can talk cheaply; cross-region adds 150ms+
Memory is ~10,000× faster than diskJustify Redis / memcached with this number

Availability Numbers#

Availability is the percentage of time a system is operational and accessible. It's often expressed as "nines."

Availability = Uptime / (Uptime + Downtime)

The nines table#

AvailabilityNinesDowntime / YearDowntime / MonthDowntime / Week
90%1 nine36.5 days72 hours16.8 hours
99%2 nines3.65 days7.2 hours1.68 hours
99.9%3 nines8.77 hours43.8 minutes10.1 minutes
99.99%4 nines52.6 minutes4.38 minutes1.01 minutes
99.999%5 nines5.26 minutes26.3 seconds6.05 seconds
99.9999%6 nines31.5 seconds2.63 seconds~1 second

Interview tip: When an interviewer asks for "high availability," they usually mean 99.9% (3 nines) or better. Pushing to 5 nines requires significant complexity — redundant components, active-active failover, zero-downtime deploys. Always ask what the SLA target is; it drives the whole architecture.

Why availability compounds#

When you chain services together, overall availability multiplies:

SetupAvailability
Two services, both 99.9%99.9% × 99.9% = 99.8%
Three services, all 99.9%99.9%³ = 99.7%
Five services, all 99.9%99.9%⁵ = 99.5%

This is why microservices need careful reliability budgets — more hops = more ways to fail.

To improve availability:

  • Redundancy — no single point of failure; replicate every critical component
  • Failover — automatic promotion of a standby (active-passive or active-active)
  • Health checks + circuit breakers — stop sending traffic to a failing node before humans notice
  • Graceful degradation — return partial results rather than a hard error

Back-of-Envelope Estimation Cheat Sheet#

Traffic#

DAUQPS (reads)Peak QPS (×3)
1 million~12~36
10 million~115~350
100 million~1,150~3,500
1 billion~11,500~35,000

Formula: QPS = DAU × (requests/user/day) ÷ 86,400 seconds

Storage#

ScenarioQuick math
10M users, each stores 1 MB/day10 TB/day
500K video uploads/day, 300 MB avg150 TB/day
100M tweets/day, 280 bytes each~28 GB/day
1B photos/day, 300 KB each300 TB/day

Rule of thumb: At 1M users storing 1 KB/day → 1 GB/day. Scale linearly from there.

Network bandwidth#

ScenarioBandwidth
1 Gbps link125 MB/s
Serving 100 MB/s of video800 Mbps
10,000 QPS, 1 KB response each~10 MB/s = 80 Mbps

Useful constants#

ValueNumber
Seconds in a day86,400 (~10^5)
Seconds in a month~2.6 million (~2.5 × 10^6)
Seconds in a year~31.5 million (~3 × 10^7)
1 Gbps =125 MB/s

Common Estimation Mistakes to Avoid#

MistakeFix
Forgetting peak vs average trafficAssume peak is 2–3× average
Ignoring replication overheadStorage × replication factor (usually 3× for distributed systems)
Confusing MB and MB/sState units explicitly every time
Treating 99% and 99.9% as "basically the same"They differ by 8 hours of downtime per year
Ignoring write amplificationWrites to replicated systems multiply: 1 write → N disk writes