NEW: ML Mock & Coaching now available

Key Concepts

High Level Design

How to create effective architecture diagrams and design scalable systems in interviews.

8 min read

The high-level design is the heart of your system design interview. This is where you translate requirements, entities, and APIs into a visual architecture that interviewers can evaluate. Done well, it demonstrates your ability to think at scale and make sound technical decisions.

The Goal of High-Level Design

Your high-level design should answer: "How do the major components work together to satisfy the requirements?"

You're not writing code. You're creating a blueprint that:

  • Shows the major components and their responsibilities
  • Illustrates how data flows through the system
  • Highlights key scaling strategies
  • Reveals your understanding of distributed systems
Info

The high-level design typically takes 15-20 minutes of a 45-minute interview. It's where most candidates are evaluated, so invest time in getting it right.

The Standard Building Blocks

Most system designs use a combination of these components:

Common Architecture Components
NameDescription
ClientsWeb, mobile, or API consumers that initiate requests
Load BalancerDistributes traffic across servers, handles SSL termination
API GatewayAuthentication, rate limiting, routing, request transformation
Application ServersStateless services that process business logic
CacheIn-memory storage for frequently accessed data (Redis, Memcached)
DatabasePersistent storage - SQL for relationships, NoSQL for flexibility
Message QueueAsync processing, decoupling services (Kafka, RabbitMQ, SQS)
CDNEdge caching for static content and reducing latency globally
Object StorageBlob storage for images, videos, files (S3, GCS)
Search IndexFull-text search capabilities (Elasticsearch, Algolia)

A Systematic Approach

Follow this framework to build your high-level design:

Step 1: Start with the Data Flow

Draw how data moves through your system for the core use case:

Client → Load Balancer → API Server → Database
                                   ↘ Cache

Step 2: Add Read/Write Paths Separately

Many systems have different read and write patterns:

Write Path (Create Tweet):

Client → LB → API Server → Database
                        → Message Queue → Async Workers (notifications, timeline fanout)

Read Path (View Timeline):

Client → LB → API Server → Cache (hit?) → Return
                                (miss?) → Database → Update Cache → Return

Step 3: Introduce Scaling Mechanisms

Based on your scale requirements, add appropriate scaling:

  • Horizontal scaling: Multiple API servers behind load balancer
  • Database sharding: Partition data across multiple databases
  • Caching layers: Reduce database load
  • CDN: Offload static content
  • Async processing: Message queues for non-critical paths

Step 4: Address Failure Scenarios

Show you've thought about reliability:

  • Database replication: Primary-replica setup
  • Multi-region: For disaster recovery
  • Circuit breakers: Graceful degradation
  • Retry mechanisms: Idempotent operations

Level-Based Expectations: Architecture

What you should emphasize varies by your target level:

High-Level Design Expectations by Level
NameDescription
Mid-Level (L4)Draw a clear, complete diagram with all major components. Explain data flow for core use cases. Show basic scaling (LB, replicas).
Senior (L5)Separate read/write paths. Discuss caching strategies. Introduce async processing. Explain database choice and sharding approach.
Staff+ (L6+)Consider multi-region architecture, consistency guarantees, failure modes, and operational concerns. Discuss evolution and tech debt.
Engineering ManagerFocus on team boundaries aligned with architecture, deployment strategies, monitoring, and how the design enables team autonomy.

Example: Twitter High-Level Design

Let's walk through a complete example:

Core Components

┌─────────────┐
│   Clients   │ (Web, iOS, Android)
└──────┬──────┘
       │
       ▼
┌─────────────┐
│     CDN     │ (Static assets, profile images)
└──────┬──────┘
       │
       ▼
┌─────────────┐
│Load Balancer│
└──────┬──────┘
       │
       ▼
┌─────────────┐
│ API Gateway │ (Auth, rate limiting)
└──────┬──────┘
       │
       ▼
┌─────────────────────────────────────┐
│         Application Services         │
│  ┌─────────┐ ┌─────────┐ ┌────────┐ │
│  │  Tweet  │ │  User   │ │Timeline│ │
│  │ Service │ │ Service │ │Service │ │
│  └────┬────┘ └────┬────┘ └────┬───┘ │
└───────┼───────────┼───────────┼─────┘
        │           │           │
        ▼           ▼           ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│Tweet Cache  │ │ User Cache  │ │Timeline     │
│  (Redis)    │ │  (Redis)    │ │Cache(Redis) │
└──────┬──────┘ └──────┬──────┘ └──────┬──────┘
       │               │               │
       ▼               ▼               ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│  Tweet DB   │ │  User DB    │ │  Graph DB   │
│ (Sharded)   │ │ (Sharded)   │ │ (Follows)   │
└─────────────┘ └─────────────┘ └─────────────┘

Async Processing

Tweet Service → Kafka → Timeline Fanout Workers → Timeline Cache Updates
                     → Notification Workers → Push to devices
                     → Search Indexer → Elasticsearch

Common Architecture Patterns

Know these patterns and when to use them:

1. Fan-out on Write vs. Fan-out on Read

Timeline Generation Strategies
NameDescription
Fan-out on WritePre-compute timelines when a tweet is posted. Fast reads, slower writes. Good for users with few followers.
Fan-out on ReadCompute timeline when requested. Slower reads, faster writes. Good for celebrities with millions of followers.
HybridFan-out on write for regular users, fan-out on read for celebrities. Twitter's actual approach.

2. CQRS (Command Query Responsibility Segregation)

Separate read and write models when their patterns differ significantly:

  • Write model: Optimized for consistency, normalized
  • Read model: Optimized for queries, denormalized
  • Event sourcing: Changes sync via events

3. Event-Driven Architecture

Decouple services with events:

Order Service → "OrderPlaced" event → Message Queue
                                          ↓
              Payment Service ← consumes event
              Inventory Service ← consumes event
              Notification Service ← consumes event
Challenge

Architecture Decision Practice

You're designing a notification system that needs to send 100 million notifications per day. How would you architect this?

Consider: What happens if sending fails? How do you handle different channels (push, email, SMS)?

See suggested architecture

Key Architecture Decisions:

  1. Message Queue as the backbone: Use Kafka for durability and high throughput

    • Producers (services) emit notification events
    • Consumers process and route to channels
  2. Channel-specific workers:

    • Push notification workers → APNs, FCM
    • Email workers → SendGrid, SES
    • SMS workers → Twilio, SNS
  3. Retry and dead-letter queues:

    • Failed notifications go to retry queue with exponential backoff
    • After N retries, move to dead-letter queue for manual inspection
  4. Rate limiting per user:

    • Prevent notification spam
    • Batch similar notifications (5 likes → "5 people liked your post")
  5. Template service:

    • Centralize notification content management
    • A/B testing for message effectiveness

Diagram:

Services → Kafka → Router → Push Queue → Push Workers → APNs/FCM
                         → Email Queue → Email Workers → SendGrid
                         → SMS Queue → SMS Workers → Twilio

           ↓ (failures)
         Retry Queue → Back to channel queues
           ↓ (max retries)
         Dead Letter Queue → Alerting

Database Selection Guide

Choose the right database for each use case:

Database Selection
NameDescription
PostgreSQL/MySQLStructured data with relationships, ACID transactions. Users, orders, payments.
MongoDBFlexible schemas, document storage. Product catalogs, content management.
RedisCaching, sessions, rate limiting, leaderboards. Sub-millisecond access.
CassandraHigh write throughput, time-series data, wide-column. IoT, logging, metrics.
ElasticsearchFull-text search, log analysis, analytics dashboards.
Neo4jGraph relationships. Social networks, recommendation engines, fraud detection.

Common Mistakes to Avoid

1. Single Points of Failure

Every critical component should have redundancy:

# Bad - single database
API Servers → Database

# Good - replicated database
API Servers → Primary DB ← replicates → Replica DB
                       ↘ Replica DB

2. Not Addressing Hot Spots

Popular content creates hot spots. Solutions:

  • Caching: Cache hot content aggressively
  • Replication: Read replicas for read-heavy hot spots
  • Sharding by access pattern: Separate hot and cold data

3. Over-Complicating the Design

Start simple, add complexity only when justified:

"I'll start with a straightforward design and then discuss optimizations for scale..."

4. Forgetting to Justify Choices

Don't just draw boxes. Explain why:

"I'm using Kafka here instead of SQS because we need exactly-once semantics and the ability to replay events..."

Warning

If you can't explain why a component is there, remove it. Interviewers notice when candidates add components they don't understand.

Presenting Your Design

Structure your walkthrough:

  1. Start with the happy path: "Here's how a user posts a tweet..."
  2. Trace the data flow: "The request goes through the load balancer to..."
  3. Highlight key decisions: "I chose Redis for caching because..."
  4. Acknowledge trade-offs: "The downside of this approach is..."
  5. Invite questions: "Does this make sense? Should I dive deeper into any component?"
Tip

Draw as you talk. A silent candidate drawing is awkward. Narrate your thought process continuously.

What's Next

Your high-level design gives the blueprint. Now it's time for deep dives, where you'll be asked to zoom into specific components and solve detailed problems.