NEW: ML Mock & Coaching now available

Key Concepts

Deep Dives

How to excel at detailed technical discussions in system design interviews.

11 min read

After presenting your high-level design, interviewers will probe specific areas in depth. These deep dives separate good candidates from great ones. They reveal whether you truly understand distributed systems or are just drawing boxes.

What Are Deep Dives?

Deep dives are focused technical discussions on specific aspects of your design. The interviewer might ask:

  • "How exactly would you implement the timeline fanout?"
  • "Walk me through what happens when this database fails"
  • "How do you ensure messages aren't lost in the queue?"
  • "What's your sharding strategy?"
Info

Deep dives typically happen in the last 10-15 minutes of the interview. The interviewer is stress-testing your design and technical knowledge.

Common Deep Dive Topics

Prepare for these areas as they come up in almost every interview:

Frequent Deep Dive Topics
NameDescription
Scaling & ShardingHow to partition data, handle hot spots, rebalance shards
CachingCache invalidation, eviction policies, cache-aside vs write-through
ConsistencyStrong vs eventual consistency, conflict resolution, consensus
Failure HandlingRetry strategies, circuit breakers, graceful degradation
Data StorageSchema design, indexing, query optimization, storage trade-offs
Rate LimitingToken bucket, sliding window, distributed rate limiting
SecurityAuthentication, authorization, encryption, data privacy
MonitoringMetrics, logging, alerting, debugging production issues

Level-Based Expectations: Deep Dives

What depth is expected at each level:

Deep Dive Expectations by Level
NameDescription
Mid-Level (L4)Explain one component in detail when asked. Understand basic concepts like caching, replication. Know when to use SQL vs NoSQL.
Senior (L5)Proactively identify areas needing deep dives. Discuss trade-offs confidently. Handle distributed systems challenges like consistency and partitioning.
Staff+ (L6+)Anticipate failure modes before asked. Discuss real-world operational experience. Connect technical decisions to business impact. Challenge assumptions.
Engineering ManagerTranslate technical complexity into team and project planning. Discuss how to incrementally build and validate. Identify risks and mitigation strategies.

Deep Dive #1: Database Sharding

One of the most common deep dive topics. Be ready to discuss:

Sharding Strategies

Sharding Approaches
NameDescription
Hash-basedDistribute by hash(key) mod N. Even distribution, but range queries are expensive.
Range-basedDistribute by key ranges. Good for range queries, but can create hot spots.
Directory-basedLookup service maps keys to shards. Flexible, but adds latency and complexity.
GeographicPartition by region. Good for latency, but cross-region operations are complex.

Key Questions to Answer

  • What's the shard key? Choose based on your access patterns
  • How do you handle hot shards? Splitting, caching, or repartitioning
  • What about cross-shard queries? Scatter-gather or denormalization
  • How do you add new shards? Consistent hashing or rebalancing strategy
Challenge

Sharding Decision

You're designing a chat application. Messages can be fetched by conversation or by user. How would you shard the messages table?

See analysis

The Challenge: Two access patterns conflict:

  1. Get all messages in a conversation (shard by conversation_id)
  2. Get all messages sent by a user (shard by user_id)

Option 1: Shard by conversation_id

  • ✅ Fast conversation reads (all messages on one shard)
  • ❌ User's messages scattered across shards

Option 2: Shard by user_id

  • ✅ Fast per-user queries
  • ❌ Group chats scattered across shards

Best Approach: Dual storage with different shard keys

  • Primary store: Sharded by conversation_id (optimized for main use case)
  • Secondary index: Sharded by user_id (for user history queries)
  • Use async replication or change data capture (CDC) to keep in sync

This is a common pattern: optimize storage for the primary access pattern, build secondary indexes for other patterns.

Deep Dive #2: Caching Strategies

Caching is critical for performance. Know these patterns:

Cache-Aside (Lazy Loading)

Read:  App checks cache → miss → read DB → populate cache → return
Write: App writes to DB → invalidate/delete cache entry

Trade-offs:

  • ✅ Only caches what's actually needed
  • ❌ First request is always slow (cache miss)
  • ❌ Cache and DB can become inconsistent

Write-Through

Write: App writes to cache → cache writes to DB → return
Read:  Always read from cache

Trade-offs:

  • ✅ Cache always consistent with DB
  • ❌ Write latency increases
  • ❌ May cache data that's never read

Write-Behind (Write-Back)

Write: App writes to cache → return immediately
       Background: cache asynchronously writes to DB

Trade-offs:

  • ✅ Fastest write performance
  • ❌ Risk of data loss if cache fails before persistence
  • ❌ Complex consistency model

Cache Invalidation

The two hard problems in computer science: cache invalidation and naming things.

Cache Invalidation Strategies
NameDescription
TTL-basedSet expiration time. Simple, but stale data until expiry.
Event-basedInvalidate on writes. Immediate consistency, but complex in distributed systems.
Version-basedInclude version in cache key. Never stale, but requires version tracking.
HybridShort TTL + event invalidation. Best of both, but more complex.

Deep Dive #3: Consistency and Availability

The CAP theorem is fundamental. Be ready to discuss:

Strong vs. Eventual Consistency

Consistency Models
NameDescription
Strong ConsistencyAll reads see the latest write. Required for: financial transactions, inventory counts. Cost: higher latency, lower availability.
Eventual ConsistencyReads may see stale data temporarily. Acceptable for: social feeds, analytics, caches. Benefit: better performance and availability.
Causal ConsistencyCausally related operations appear in order. Good middle ground for many applications.
Read-your-writesUsers see their own writes immediately. Critical for good user experience.

Handling Conflicts

In eventually consistent systems, conflicts happen. Resolution strategies:

  • Last-write-wins (LWW): Simple but can lose data
  • Vector clocks: Track causality, detect conflicts
  • Application-level merge: Custom logic (e.g., CRDTs for collaborative editing)
Challenge

Consistency Trade-off

You're designing a shopping cart service. What consistency model would you choose and why?

See analysis

The key insight: Different operations need different consistency.

Adding items to cart: Eventual consistency is fine

  • Users tolerate brief delays in seeing items added
  • Performance and availability are more important
  • Conflict resolution: merge cart items (union)

Checkout/payment: Strong consistency required

  • Must accurately reflect current cart contents
  • Must check inventory availability
  • Can't lose or duplicate orders

Recommended approach:

  1. Cart storage: Eventually consistent, replicated across regions
  2. Checkout flow:
    • Read cart with strong consistency (read from primary)
    • Use distributed transaction or saga pattern for payment
    • Reserve inventory before charging

This hybrid approach gives fast performance for browsing while ensuring correctness for transactions.

Deep Dive #4: Failure Handling

Distributed systems fail. Show you've thought about this:

Retry Strategies

# Exponential backoff with jitter
retry_delay = min(base_delay * (2 ^ attempt) + random_jitter, max_delay)

Key considerations:

  • Idempotency: Is it safe to retry?
  • Timeout: How long to wait?
  • Max retries: When to give up?
  • Circuit breaker: When to stop trying altogether?

Circuit Breaker Pattern

Prevent cascade failures:

State: CLOSED → OPEN → HALF-OPEN → CLOSED

CLOSED:  Requests flow normally, failures counted
         If failures exceed threshold → OPEN

OPEN:    Requests fail immediately (fast failure)
         After timeout → HALF-OPEN

HALF-OPEN: Allow one request through
           If succeeds → CLOSED
           If fails → OPEN

Graceful Degradation

When a dependency fails, degrade gracefully:

Degradation Strategies
NameDescription
Serve stale dataReturn cached data even if expired. Better than error.
Reduce functionalityDisable non-critical features. Keep core working.
Queue for laterAccept writes, process when service recovers.
Provide estimatesShow approximate data when exact is unavailable.

Deep Dive #5: Rate Limiting

Protect your system from abuse and overload:

Algorithms

Rate Limiting Algorithms
NameDescription
Token BucketTokens added at fixed rate, consumed per request. Allows bursts. Most common.
Leaky BucketRequests queue and process at fixed rate. Smooths traffic. Good for APIs.
Fixed WindowCount requests per time window. Simple but allows burst at window boundary.
Sliding Window LogTrack timestamp of each request. Accurate but memory intensive.
Sliding Window CounterWeighted count between windows. Good balance of accuracy and efficiency.

Distributed Rate Limiting

Challenge: Rate limits across multiple servers.

Approaches:

  1. Centralized counter (Redis): Accurate but adds latency
  2. Local counters with sync: Fast but can over-allow
  3. Sticky sessions: Route user to same server, local counting
Challenge

Rate Limiting Design

Design a rate limiter for an API that allows 100 requests per minute per user, across 10 API servers.

See implementation approach

Architecture:

Client → Load Balancer → API Server → Redis (rate limit check) → Backend

Implementation with Token Bucket in Redis:

def is_allowed(user_id: str) -> bool:
    key = f"rate_limit:{user_id}"
    now = time.time()

    # Use Redis transaction
    with redis.pipeline() as pipe:
        # Get current bucket state
        pipe.hgetall(key)
        # Execute and get result
        result = pipe.execute()[0]

        tokens = float(result.get('tokens', 100))
        last_update = float(result.get('last_update', now))

        # Refill tokens based on time elapsed
        elapsed = now - last_update
        tokens = min(100, tokens + elapsed * (100/60))  # 100 tokens per 60 sec

        if tokens >= 1:
            # Consume token
            tokens -= 1
            pipe.hset(key, mapping={'tokens': tokens, 'last_update': now})
            pipe.expire(key, 120)  # Cleanup
            pipe.execute()
            return True
        else:
            return False

Key points:

  • Use Redis for distributed state (atomic operations)
  • Token bucket allows short bursts while maintaining average rate
  • Set TTL on keys for automatic cleanup
  • Consider using Redis Lua scripts for true atomicity

Optimizations for high scale:

  • Local cache with async sync to reduce Redis calls
  • Probabilistic rate limiting (approximate counting)
  • Separate rate limit service to isolate failures

How to Handle Deep Dives You Don't Know

You won't know everything. Here's how to handle gaps:

Be Honest

"I haven't implemented this specific pattern before, but here's how I'd approach it..."

Reason from First Principles

"I'm not sure of the exact algorithm, but the problem we're solving is X, so the solution likely needs to handle Y and Z..."

Ask Clarifying Questions

"To make sure I'm focusing on the right aspect, are you more interested in the consistency guarantees or the performance characteristics?"

Tip

Interviewers respect intellectual honesty. They're evaluating how you think, not testing memorization. A thoughtful "I don't know, but..." is better than a wrong confident answer.

Common Deep Dive Mistakes

1. Going Too Deep Too Early

Don't spend 10 minutes on one component before finishing the high-level design.

2. Not Adapting to Feedback

If the interviewer looks confused or tries to redirect, pivot quickly.

3. Memorized Answers Without Understanding

Interviewers can tell when you've memorized something. Be ready for follow-up "why" questions.

4. Ignoring Trade-offs

Every technical decision has trade-offs. Acknowledge them proactively:

"This approach gives us strong consistency but increases latency. Given our requirements, I think that's the right trade-off because..."

Putting It All Together

A successful deep dive looks like this:

Interviewer: "How would you handle the case where a celebrity with 10 million followers posts a tweet?"

You: "Great question. This is the hot follower problem. If we fan out on write, a single tweet would trigger 10 million timeline updates, which would be too slow.

I'd use a hybrid approach: for regular users, we continue fan-out on write for fast reads. For celebrities, we mark them as 'high-follower' accounts and use fan-out on read instead.

When a user fetches their timeline, we merge two sources: their pre-computed timeline cache, plus a live query of recent tweets from any celebrities they follow.

The trade-off is more complex read path, but it avoids the write amplification problem. We'd monitor the threshold for what counts as 'high-follower' and tune it based on system performance."

Info

Notice the structure: identify the problem, propose solution, explain how it works, acknowledge trade-offs, mention operational considerations.

What's Next

You now have a framework for requirements, entities/APIs, high-level design, and deep dives. The next step is practice. Work through real problems, get feedback, and iterate.