Caching is one of the most impactful optimizations in system design. A well-designed cache can reduce latency from hundreds of milliseconds to single-digit milliseconds and dramatically reduce database load.
Why Caching Matters
Without caching, every request hits your database. At scale, this creates:
- High latency: Database queries take 10-100ms
- Database overload: Thousands of queries per second
- Poor user experience: Slow page loads
- High costs: Over-provisioned databases
With caching, frequently accessed data is served from memory in under 1ms.
The rule of thumb: if data is read more than it's written, it's a candidate for caching.
Caching Technologies
| Popular Caching Solutions | |
|---|---|
| Name | Description |
Redis | In-memory data store with rich data structures (strings, lists, sets, sorted sets, hashes). Supports persistence, pub/sub, and Lua scripting. |
Memcached | Simple, high-performance key-value cache. No persistence, limited data types. Great for simple caching needs. |
Application Cache | In-process caching (e.g., Guava, Caffeine). Fastest access but limited to single instance. Good for static configuration. |
CDN Cache | Edge caching for static assets. Reduces latency globally. Covered in the CDN section. |
Redis vs Memcached
| Redis vs Memcached Comparison | |
|---|---|
| Name | Description |
Data structures | Redis: Rich (lists, sets, sorted sets, hashes). Memcached: Strings only. |
Persistence | Redis: Optional (RDB, AOF). Memcached: None. |
Clustering | Redis: Built-in cluster mode. Memcached: Client-side sharding. |
Memory efficiency | Redis: Less efficient. Memcached: More efficient for simple strings. |
Use case | Redis: Complex caching, leaderboards, sessions. Memcached: Simple key-value caching at scale. |
Interview default: Use Redis unless you have a specific reason for Memcached. Redis's flexibility makes it suitable for most scenarios.
Caching Strategies
Cache-Aside (Lazy Loading)
The most common pattern. Application manages the cache explicitly.
Read:
1. Check cache for data
2. If hit → return cached data
3. If miss → query database → store in cache → return
Write:
1. Write to database
2. Invalidate or update cache
Pros: Only caches what's needed, simple to implement Cons: First request always slow, potential for stale data
Read-Through
Cache sits between application and database. Cache handles misses automatically.
Read:
1. Application requests from cache
2. Cache checks if data exists
3. If miss → cache queries database → stores result → returns
Pros: Application code is simpler Cons: Requires cache that supports read-through
Write-Through
Writes go through the cache to the database.
Write:
1. Application writes to cache
2. Cache synchronously writes to database
3. Returns success when both complete
Pros: Cache always consistent with database Cons: Higher write latency, may cache data that's never read
Write-Behind (Write-Back)
Cache buffers writes and asynchronously persists to database.
Write:
1. Application writes to cache
2. Cache returns success immediately
3. Cache asynchronously batches writes to database
Pros: Fastest write performance, batching reduces database load Cons: Risk of data loss if cache fails before persisting
Choose the Right Strategy
You're designing a social media feed. Users see their timeline frequently (reads), and new posts are added occasionally (writes). Which caching strategy would you use?
See recommendation
Recommended: Cache-Aside with TTL
- Why cache-aside: Feed data is read-heavy, and we only want to cache active users' feeds
- Why TTL: Feeds change over time; a 5-minute TTL ensures reasonable freshness without constant invalidation
- Write handling: When a user posts, invalidate their followers' cached feeds (or let TTL handle it)
Alternative consideration: For celebrity accounts with millions of followers, fan-out on write (pre-computing feeds) combined with caching avoids the thundering herd problem when their cache expires.
Cache Invalidation
The hardest problem in caching. How do you ensure cached data stays consistent with the source of truth?
TTL-Based Expiration
Set a time-to-live on cached entries.
SET user:123:profile "{...}" EX 3600 # Expires in 1 hour
Pros: Simple, automatic cleanup Cons: Data can be stale until expiration
Event-Based Invalidation
Invalidate cache when data changes.
# On user profile update:
UPDATE users SET name = 'New Name' WHERE id = 123;
DELETE FROM cache WHERE key = 'user:123:profile';
Pros: Always fresh data Cons: Complex in distributed systems, easy to miss invalidation
Version-Based Keys
Include version in cache key.
user:123:profile:v5 # Version 5 of user's profile
Pros: Never serve stale data, supports rollback Cons: Need to track versions, old entries accumulate
Cache invalidation bugs are subtle and hard to debug. When in doubt, prefer shorter TTLs over complex invalidation logic.
Common Caching Patterns
Cache Warming
Pre-populate cache before traffic arrives.
# On deployment or startup:
for user in get_active_users():
cache.set(f"user:{user.id}:profile", user.to_json())
Use when: You know what data will be accessed, cold cache would cause problems.
Thundering Herd Protection
When a popular cache entry expires, many requests hit the database simultaneously.
Solutions:
- Locking: Only one request fetches from DB, others wait
- Probabilistic expiration: Randomly refresh before TTL
- Background refresh: Refresh cache before expiration
# Probabilistic early expiration
ttl_remaining = cache.ttl(key)
if ttl_remaining < 60 and random.random() < 0.1: # 10% chance
refresh_cache_async(key)
Multi-Level Caching
Layer caches for optimal performance.
L1: Application memory (microseconds)
L2: Redis/Memcached (milliseconds)
L3: Database (tens of milliseconds)
Use when: Extremely high throughput requirements, hot data fits in memory.
Level-Based Expectations
| Caching Knowledge by Level | |
|---|---|
| Name | Description |
Mid-Level (L4) | Know when to add a cache layer. Understand TTL-based expiration. Can explain cache-aside pattern. |
Senior (L5) | Discuss cache invalidation strategies and trade-offs. Handle thundering herd. Choose between Redis/Memcached with reasoning. |
Staff+ (L6+) | Design multi-level caching architectures. Discuss operational concerns (monitoring, memory pressure, failover). Consider cache consistency in distributed systems. |
Engineering Manager | Evaluate caching infrastructure costs and trade-offs. Understand team implications of cache complexity. Plan for cache-related incidents. |
Common Interview Mistakes
1. Caching Everything
Not all data benefits from caching. Consider:
- Is it read more than written?
- Is the data accessed frequently?
- Can you tolerate some staleness?
2. Ignoring Cache Failures
What happens when Redis is down? Your system should degrade gracefully, not crash.
try:
data = cache.get(key)
except CacheError:
data = None # Fall back to database
if data is None:
data = database.get(key)
3. Not Discussing Eviction Policies
When cache is full, what gets evicted?
| Cache Eviction Policies | |
|---|---|
| Name | Description |
LRU (Least Recently Used) | Evict data that hasn't been accessed recently. Most common, good default. |
LFU (Least Frequently Used) | Evict data accessed least often. Better for some workloads but more complex. |
FIFO (First In First Out) | Evict oldest data. Simple but not access-aware. |
Random | Randomly evict. Surprisingly effective, very simple. |
What's Next
Caching handles read performance. Next, we'll look at load balancers, which distribute traffic across your servers.