High Level Design
Caching
How caching works, the five write strategies (read-through, cache-aside, write-through, write-around, write-back), eviction policies, and common pitfalls.
Caching is storing frequently accessed data in temporary storage to improve both latency (time per operation) and throughput (operation rate).
Two key parameters:
- Write Policy — how to sync writes between cache and DB
- Eviction / Replacement Policy — what to kick out when the cache is full (LRU, LFU, …)
Why Caching?#
- Reduce network calls
- Avoid repeated database queries
- Absorb uneven loads and traffic spikes
- Improve page load times
How Caching Works (Cache-Aside Flow)#
- Look for entry in cache → cache miss
- Load entry from the database
- Add entry to cache
- Return entry to the caller
Drawbacks#
- Poor hit rate if cache doesn't store what users need
- Must maintain consistency between cache and DB (cache invalidation → eventual consistency)
- Cache Thrashing — constantly evicting and reloading the same items
Types of Caches#
| Type | Description | Examples |
|---|---|---|
| In-Memory Cache | Data in RAM; extremely fast access | Redis, Memcached |
| Distributed Cache | Spans multiple servers; for large-scale systems | Redis Cluster, Amazon ElastiCache |
| Client-Side Cache | Stored on client device (cookies, local storage, browser cache) | Browser cache |
| Database Cache | Stores frequently queried results; reduces DB queries | Query result caches |
| CDN | Stores content on geographically distributed servers | Cloudflare, CloudFront |
Caching Strategies#
Read-Through#
The cache acts as an intermediary. On a miss, the cache itself fetches from the DB and populates itself.
- ✅ Simplifies application logic
- ✅ Only frequently accessed data stays cached
- ❌ Initial read has extra latency (cache fetches from DB)
- Use TTL to avoid stale data

Best for: Read-heavy apps where data is accessed often but updated rarely — CDNs, social media feeds, user profiles.
Cache-Aside (Lazy Loading)#
The application handles reading from cache and loading on misses. Cache is a "sidecar" to the database.
- Check cache → hit: return data; miss: fetch from DB
- Load fetched data into cache for next request
- Set TTL to prevent stale data

Best for: High read-to-write ratio with infrequent updates — e.g., e-commerce product catalog.
Write-Through#
Every write goes to both the cache and DB simultaneously. Synchronous — no delay in data propagation.
- ✅ Strong consistency; cache always has latest data
- ✅ No TTL needed in most cases
- ❌ Higher write latency (writes to two places)

Best for: Consistency-critical systems — financial apps, online transaction processing.
Write-Around#
Writes go directly to the DB, bypassing the cache. Cache is populated only on subsequent reads (via cache-aside).
- ✅ Cache stays clean — only frequently accessed data gets cached
- ✅ Writes are faster (single destination)
- ❌ First read after write always hits the DB (cache miss)

Best for: Write-heavy systems where data is rarely re-read — logging, audit trails.
Write-Back (Write-Behind)#
Data is written to cache first, then asynchronously written to DB in the background.
- ✅ Lowest write latency
- ❌ Risk of data loss if cache fails before syncing to DB
- Mitigated with Redis AOF (Append Only File) for persistence

Best for: Write-heavy scenarios where speed matters over immediate consistency — gaming state, social media feeds.
Write Policy Comparison#
| Feature | Write-Through | Write-Back | Write-Around |
|---|---|---|---|
| Write Location | Cache ✅ + DB ✅ | Cache ✅ only (DB later) | DB ✅ only |
| Read After Write | Fast ✅ (in cache) | Fast ✅ (in cache) | Slow ❌ (not in cache yet) |
| Write Latency | Slower ❌ | Fast ✅ | Fast ✅ |
| Cache Pollution | Possible ❌ | Possible ❌ | Avoided ✅ |
| Data Consistency | High ✅ | Lower ❌ (risk if cache lost) | High ✅ |

Quick Use-Case Reference#
| Scenario | Strategy |
|---|---|
| Financial transactions, logs | Write-Through |
| Gaming state or temporary calculations | Write-Back |
| Logging, rarely-read audit trails | Write-Around |
Cache Eviction Policies#
When the cache is full, these policies decide what to remove:
| Policy | Description |
|---|---|
| LRU (Least Recently Used) | Evicts the least recently accessed item |
| LFU (Least Frequently Used) | Evicts items accessed the fewest times |
| FIFO | Evicts the oldest item regardless of access pattern |
| TTL (Time-to-Live) | Removes items after a fixed duration |
| MRU (Most Recently Used) | Removes the most recently used item |
| ARC (Adaptive Replacement Cache) | Balances between LRU and LFU dynamically |
Memcached uses Segmented LRU (blend of LRU + LFU).
Challenges and Considerations#
| Challenge | Description |
|---|---|
| Cache Coherence | Keeping cache consistent with the source of truth (DB) |
| Cache Invalidation | Knowing when and how to remove stale data |
| Cold Start | Handling an empty cache after system restart |
| Cache Penetration | Repeated queries for non-existent data overwhelm the backend |
| Cache Stampede | Many concurrent requests try to rebuild the cache simultaneously |