After presenting your high-level design, interviewers will probe specific areas in depth. These deep dives separate good candidates from great ones. They reveal whether you truly understand distributed systems or are just drawing boxes.
What Are Deep Dives?
Deep dives are focused technical discussions on specific aspects of your design. The interviewer might ask:
- "How exactly would you implement the timeline fanout?"
- "Walk me through what happens when this database fails"
- "How do you ensure messages aren't lost in the queue?"
- "What's your sharding strategy?"
Deep dives typically happen in the last 10-15 minutes of the interview. The interviewer is stress-testing your design and technical knowledge.
Common Deep Dive Topics
Prepare for these areas as they come up in almost every interview:
| Frequent Deep Dive Topics | |
|---|---|
| Name | Description |
Scaling & Sharding | How to partition data, handle hot spots, rebalance shards |
Caching | Cache invalidation, eviction policies, cache-aside vs write-through |
Consistency | Strong vs eventual consistency, conflict resolution, consensus |
Failure Handling | Retry strategies, circuit breakers, graceful degradation |
Data Storage | Schema design, indexing, query optimization, storage trade-offs |
Rate Limiting | Token bucket, sliding window, distributed rate limiting |
Security | Authentication, authorization, encryption, data privacy |
Monitoring | Metrics, logging, alerting, debugging production issues |
Level-Based Expectations: Deep Dives
What depth is expected at each level:
| Deep Dive Expectations by Level | |
|---|---|
| Name | Description |
Mid-Level (L4) | Explain one component in detail when asked. Understand basic concepts like caching, replication. Know when to use SQL vs NoSQL. |
Senior (L5) | Proactively identify areas needing deep dives. Discuss trade-offs confidently. Handle distributed systems challenges like consistency and partitioning. |
Staff+ (L6+) | Anticipate failure modes before asked. Discuss real-world operational experience. Connect technical decisions to business impact. Challenge assumptions. |
Engineering Manager | Translate technical complexity into team and project planning. Discuss how to incrementally build and validate. Identify risks and mitigation strategies. |
Deep Dive #1: Database Sharding
One of the most common deep dive topics. Be ready to discuss:
Sharding Strategies
| Sharding Approaches | |
|---|---|
| Name | Description |
Hash-based | Distribute by hash(key) mod N. Even distribution, but range queries are expensive. |
Range-based | Distribute by key ranges. Good for range queries, but can create hot spots. |
Directory-based | Lookup service maps keys to shards. Flexible, but adds latency and complexity. |
Geographic | Partition by region. Good for latency, but cross-region operations are complex. |
Key Questions to Answer
- What's the shard key? Choose based on your access patterns
- How do you handle hot shards? Splitting, caching, or repartitioning
- What about cross-shard queries? Scatter-gather or denormalization
- How do you add new shards? Consistent hashing or rebalancing strategy
Sharding Decision
You're designing a chat application. Messages can be fetched by conversation or by user. How would you shard the messages table?
See analysis
The Challenge: Two access patterns conflict:
- Get all messages in a conversation (shard by conversation_id)
- Get all messages sent by a user (shard by user_id)
Option 1: Shard by conversation_id
- ✅ Fast conversation reads (all messages on one shard)
- ❌ User's messages scattered across shards
Option 2: Shard by user_id
- ✅ Fast per-user queries
- ❌ Group chats scattered across shards
Best Approach: Dual storage with different shard keys
- Primary store: Sharded by conversation_id (optimized for main use case)
- Secondary index: Sharded by user_id (for user history queries)
- Use async replication or change data capture (CDC) to keep in sync
This is a common pattern: optimize storage for the primary access pattern, build secondary indexes for other patterns.
Deep Dive #2: Caching Strategies
Caching is critical for performance. Know these patterns:
Cache-Aside (Lazy Loading)
Read: App checks cache → miss → read DB → populate cache → return
Write: App writes to DB → invalidate/delete cache entry
Trade-offs:
- ✅ Only caches what's actually needed
- ❌ First request is always slow (cache miss)
- ❌ Cache and DB can become inconsistent
Write-Through
Write: App writes to cache → cache writes to DB → return
Read: Always read from cache
Trade-offs:
- ✅ Cache always consistent with DB
- ❌ Write latency increases
- ❌ May cache data that's never read
Write-Behind (Write-Back)
Write: App writes to cache → return immediately
Background: cache asynchronously writes to DB
Trade-offs:
- ✅ Fastest write performance
- ❌ Risk of data loss if cache fails before persistence
- ❌ Complex consistency model
Cache Invalidation
The two hard problems in computer science: cache invalidation and naming things.
| Cache Invalidation Strategies | |
|---|---|
| Name | Description |
TTL-based | Set expiration time. Simple, but stale data until expiry. |
Event-based | Invalidate on writes. Immediate consistency, but complex in distributed systems. |
Version-based | Include version in cache key. Never stale, but requires version tracking. |
Hybrid | Short TTL + event invalidation. Best of both, but more complex. |
Deep Dive #3: Consistency and Availability
The CAP theorem is fundamental. Be ready to discuss:
Strong vs. Eventual Consistency
| Consistency Models | |
|---|---|
| Name | Description |
Strong Consistency | All reads see the latest write. Required for: financial transactions, inventory counts. Cost: higher latency, lower availability. |
Eventual Consistency | Reads may see stale data temporarily. Acceptable for: social feeds, analytics, caches. Benefit: better performance and availability. |
Causal Consistency | Causally related operations appear in order. Good middle ground for many applications. |
Read-your-writes | Users see their own writes immediately. Critical for good user experience. |
Handling Conflicts
In eventually consistent systems, conflicts happen. Resolution strategies:
- Last-write-wins (LWW): Simple but can lose data
- Vector clocks: Track causality, detect conflicts
- Application-level merge: Custom logic (e.g., CRDTs for collaborative editing)
Consistency Trade-off
You're designing a shopping cart service. What consistency model would you choose and why?
See analysis
The key insight: Different operations need different consistency.
Adding items to cart: Eventual consistency is fine
- Users tolerate brief delays in seeing items added
- Performance and availability are more important
- Conflict resolution: merge cart items (union)
Checkout/payment: Strong consistency required
- Must accurately reflect current cart contents
- Must check inventory availability
- Can't lose or duplicate orders
Recommended approach:
- Cart storage: Eventually consistent, replicated across regions
- Checkout flow:
- Read cart with strong consistency (read from primary)
- Use distributed transaction or saga pattern for payment
- Reserve inventory before charging
This hybrid approach gives fast performance for browsing while ensuring correctness for transactions.
Deep Dive #4: Failure Handling
Distributed systems fail. Show you've thought about this:
Retry Strategies
# Exponential backoff with jitter
retry_delay = min(base_delay * (2 ^ attempt) + random_jitter, max_delay)
Key considerations:
- Idempotency: Is it safe to retry?
- Timeout: How long to wait?
- Max retries: When to give up?
- Circuit breaker: When to stop trying altogether?
Circuit Breaker Pattern
Prevent cascade failures:
State: CLOSED → OPEN → HALF-OPEN → CLOSED
CLOSED: Requests flow normally, failures counted
If failures exceed threshold → OPEN
OPEN: Requests fail immediately (fast failure)
After timeout → HALF-OPEN
HALF-OPEN: Allow one request through
If succeeds → CLOSED
If fails → OPEN
Graceful Degradation
When a dependency fails, degrade gracefully:
| Degradation Strategies | |
|---|---|
| Name | Description |
Serve stale data | Return cached data even if expired. Better than error. |
Reduce functionality | Disable non-critical features. Keep core working. |
Queue for later | Accept writes, process when service recovers. |
Provide estimates | Show approximate data when exact is unavailable. |
Deep Dive #5: Rate Limiting
Protect your system from abuse and overload:
Algorithms
| Rate Limiting Algorithms | |
|---|---|
| Name | Description |
Token Bucket | Tokens added at fixed rate, consumed per request. Allows bursts. Most common. |
Leaky Bucket | Requests queue and process at fixed rate. Smooths traffic. Good for APIs. |
Fixed Window | Count requests per time window. Simple but allows burst at window boundary. |
Sliding Window Log | Track timestamp of each request. Accurate but memory intensive. |
Sliding Window Counter | Weighted count between windows. Good balance of accuracy and efficiency. |
Distributed Rate Limiting
Challenge: Rate limits across multiple servers.
Approaches:
- Centralized counter (Redis): Accurate but adds latency
- Local counters with sync: Fast but can over-allow
- Sticky sessions: Route user to same server, local counting
Rate Limiting Design
Design a rate limiter for an API that allows 100 requests per minute per user, across 10 API servers.
See implementation approach
Architecture:
Client → Load Balancer → API Server → Redis (rate limit check) → Backend
Implementation with Token Bucket in Redis:
def is_allowed(user_id: str) -> bool:
key = f"rate_limit:{user_id}"
now = time.time()
# Use Redis transaction
with redis.pipeline() as pipe:
# Get current bucket state
pipe.hgetall(key)
# Execute and get result
result = pipe.execute()[0]
tokens = float(result.get('tokens', 100))
last_update = float(result.get('last_update', now))
# Refill tokens based on time elapsed
elapsed = now - last_update
tokens = min(100, tokens + elapsed * (100/60)) # 100 tokens per 60 sec
if tokens >= 1:
# Consume token
tokens -= 1
pipe.hset(key, mapping={'tokens': tokens, 'last_update': now})
pipe.expire(key, 120) # Cleanup
pipe.execute()
return True
else:
return False
Key points:
- Use Redis for distributed state (atomic operations)
- Token bucket allows short bursts while maintaining average rate
- Set TTL on keys for automatic cleanup
- Consider using Redis Lua scripts for true atomicity
Optimizations for high scale:
- Local cache with async sync to reduce Redis calls
- Probabilistic rate limiting (approximate counting)
- Separate rate limit service to isolate failures
How to Handle Deep Dives You Don't Know
You won't know everything. Here's how to handle gaps:
Be Honest
"I haven't implemented this specific pattern before, but here's how I'd approach it..."
Reason from First Principles
"I'm not sure of the exact algorithm, but the problem we're solving is X, so the solution likely needs to handle Y and Z..."
Ask Clarifying Questions
"To make sure I'm focusing on the right aspect, are you more interested in the consistency guarantees or the performance characteristics?"
Interviewers respect intellectual honesty. They're evaluating how you think, not testing memorization. A thoughtful "I don't know, but..." is better than a wrong confident answer.
Common Deep Dive Mistakes
1. Going Too Deep Too Early
Don't spend 10 minutes on one component before finishing the high-level design.
2. Not Adapting to Feedback
If the interviewer looks confused or tries to redirect, pivot quickly.
3. Memorized Answers Without Understanding
Interviewers can tell when you've memorized something. Be ready for follow-up "why" questions.
4. Ignoring Trade-offs
Every technical decision has trade-offs. Acknowledge them proactively:
"This approach gives us strong consistency but increases latency. Given our requirements, I think that's the right trade-off because..."
Putting It All Together
A successful deep dive looks like this:
Interviewer: "How would you handle the case where a celebrity with 10 million followers posts a tweet?"
You: "Great question. This is the hot follower problem. If we fan out on write, a single tweet would trigger 10 million timeline updates, which would be too slow.
I'd use a hybrid approach: for regular users, we continue fan-out on write for fast reads. For celebrities, we mark them as 'high-follower' accounts and use fan-out on read instead.
When a user fetches their timeline, we merge two sources: their pre-computed timeline cache, plus a live query of recent tweets from any celebrities they follow.
The trade-off is more complex read path, but it avoids the write amplification problem. We'd monitor the threshold for what counts as 'high-follower' and tune it based on system performance."
Notice the structure: identify the problem, propose solution, explain how it works, acknowledge trade-offs, mention operational considerations.
What's Next
You now have a framework for requirements, entities/APIs, high-level design, and deep dives. The next step is practice. Work through real problems, get feedback, and iterate.