The high-level design is the heart of your system design interview. This is where you translate requirements, entities, and APIs into a visual architecture that interviewers can evaluate. Done well, it demonstrates your ability to think at scale and make sound technical decisions.
The Goal of High-Level Design
Your high-level design should answer: "How do the major components work together to satisfy the requirements?"
You're not writing code. You're creating a blueprint that:
- Shows the major components and their responsibilities
- Illustrates how data flows through the system
- Highlights key scaling strategies
- Reveals your understanding of distributed systems
The high-level design typically takes 15-20 minutes of a 45-minute interview. It's where most candidates are evaluated, so invest time in getting it right.
The Standard Building Blocks
Most system designs use a combination of these components:
| Common Architecture Components | |
|---|---|
| Name | Description |
Clients | Web, mobile, or API consumers that initiate requests |
Load Balancer | Distributes traffic across servers, handles SSL termination |
API Gateway | Authentication, rate limiting, routing, request transformation |
Application Servers | Stateless services that process business logic |
Cache | In-memory storage for frequently accessed data (Redis, Memcached) |
Database | Persistent storage - SQL for relationships, NoSQL for flexibility |
Message Queue | Async processing, decoupling services (Kafka, RabbitMQ, SQS) |
CDN | Edge caching for static content and reducing latency globally |
Object Storage | Blob storage for images, videos, files (S3, GCS) |
Search Index | Full-text search capabilities (Elasticsearch, Algolia) |
A Systematic Approach
Follow this framework to build your high-level design:
Step 1: Start with the Data Flow
Draw how data moves through your system for the core use case:
Client → Load Balancer → API Server → Database
↘ Cache
Step 2: Add Read/Write Paths Separately
Many systems have different read and write patterns:
Write Path (Create Tweet):
Client → LB → API Server → Database
→ Message Queue → Async Workers (notifications, timeline fanout)
Read Path (View Timeline):
Client → LB → API Server → Cache (hit?) → Return
(miss?) → Database → Update Cache → Return
Step 3: Introduce Scaling Mechanisms
Based on your scale requirements, add appropriate scaling:
- Horizontal scaling: Multiple API servers behind load balancer
- Database sharding: Partition data across multiple databases
- Caching layers: Reduce database load
- CDN: Offload static content
- Async processing: Message queues for non-critical paths
Step 4: Address Failure Scenarios
Show you've thought about reliability:
- Database replication: Primary-replica setup
- Multi-region: For disaster recovery
- Circuit breakers: Graceful degradation
- Retry mechanisms: Idempotent operations
Level-Based Expectations: Architecture
What you should emphasize varies by your target level:
| High-Level Design Expectations by Level | |
|---|---|
| Name | Description |
Mid-Level (L4) | Draw a clear, complete diagram with all major components. Explain data flow for core use cases. Show basic scaling (LB, replicas). |
Senior (L5) | Separate read/write paths. Discuss caching strategies. Introduce async processing. Explain database choice and sharding approach. |
Staff+ (L6+) | Consider multi-region architecture, consistency guarantees, failure modes, and operational concerns. Discuss evolution and tech debt. |
Engineering Manager | Focus on team boundaries aligned with architecture, deployment strategies, monitoring, and how the design enables team autonomy. |
Example: Twitter High-Level Design
Let's walk through a complete example:
Core Components
┌─────────────┐
│ Clients │ (Web, iOS, Android)
└──────┬──────┘
│
▼
┌─────────────┐
│ CDN │ (Static assets, profile images)
└──────┬──────┘
│
▼
┌─────────────┐
│Load Balancer│
└──────┬──────┘
│
▼
┌─────────────┐
│ API Gateway │ (Auth, rate limiting)
└──────┬──────┘
│
▼
┌─────────────────────────────────────┐
│ Application Services │
│ ┌─────────┐ ┌─────────┐ ┌────────┐ │
│ │ Tweet │ │ User │ │Timeline│ │
│ │ Service │ │ Service │ │Service │ │
│ └────┬────┘ └────┬────┘ └────┬───┘ │
└───────┼───────────┼───────────┼─────┘
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│Tweet Cache │ │ User Cache │ │Timeline │
│ (Redis) │ │ (Redis) │ │Cache(Redis) │
└──────┬──────┘ └──────┬──────┘ └──────┬──────┘
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Tweet DB │ │ User DB │ │ Graph DB │
│ (Sharded) │ │ (Sharded) │ │ (Follows) │
└─────────────┘ └─────────────┘ └─────────────┘
Async Processing
Tweet Service → Kafka → Timeline Fanout Workers → Timeline Cache Updates
→ Notification Workers → Push to devices
→ Search Indexer → Elasticsearch
Common Architecture Patterns
Know these patterns and when to use them:
1. Fan-out on Write vs. Fan-out on Read
| Timeline Generation Strategies | |
|---|---|
| Name | Description |
Fan-out on Write | Pre-compute timelines when a tweet is posted. Fast reads, slower writes. Good for users with few followers. |
Fan-out on Read | Compute timeline when requested. Slower reads, faster writes. Good for celebrities with millions of followers. |
Hybrid | Fan-out on write for regular users, fan-out on read for celebrities. Twitter's actual approach. |
2. CQRS (Command Query Responsibility Segregation)
Separate read and write models when their patterns differ significantly:
- Write model: Optimized for consistency, normalized
- Read model: Optimized for queries, denormalized
- Event sourcing: Changes sync via events
3. Event-Driven Architecture
Decouple services with events:
Order Service → "OrderPlaced" event → Message Queue
↓
Payment Service ← consumes event
Inventory Service ← consumes event
Notification Service ← consumes event
Architecture Decision Practice
You're designing a notification system that needs to send 100 million notifications per day. How would you architect this?
Consider: What happens if sending fails? How do you handle different channels (push, email, SMS)?
See suggested architecture
Key Architecture Decisions:
-
Message Queue as the backbone: Use Kafka for durability and high throughput
- Producers (services) emit notification events
- Consumers process and route to channels
-
Channel-specific workers:
- Push notification workers → APNs, FCM
- Email workers → SendGrid, SES
- SMS workers → Twilio, SNS
-
Retry and dead-letter queues:
- Failed notifications go to retry queue with exponential backoff
- After N retries, move to dead-letter queue for manual inspection
-
Rate limiting per user:
- Prevent notification spam
- Batch similar notifications (5 likes → "5 people liked your post")
-
Template service:
- Centralize notification content management
- A/B testing for message effectiveness
Diagram:
Services → Kafka → Router → Push Queue → Push Workers → APNs/FCM
→ Email Queue → Email Workers → SendGrid
→ SMS Queue → SMS Workers → Twilio
↓ (failures)
Retry Queue → Back to channel queues
↓ (max retries)
Dead Letter Queue → Alerting
Database Selection Guide
Choose the right database for each use case:
| Database Selection | |
|---|---|
| Name | Description |
PostgreSQL/MySQL | Structured data with relationships, ACID transactions. Users, orders, payments. |
MongoDB | Flexible schemas, document storage. Product catalogs, content management. |
Redis | Caching, sessions, rate limiting, leaderboards. Sub-millisecond access. |
Cassandra | High write throughput, time-series data, wide-column. IoT, logging, metrics. |
Elasticsearch | Full-text search, log analysis, analytics dashboards. |
Neo4j | Graph relationships. Social networks, recommendation engines, fraud detection. |
Common Mistakes to Avoid
1. Single Points of Failure
Every critical component should have redundancy:
# Bad - single database
API Servers → Database
# Good - replicated database
API Servers → Primary DB ← replicates → Replica DB
↘ Replica DB
2. Not Addressing Hot Spots
Popular content creates hot spots. Solutions:
- Caching: Cache hot content aggressively
- Replication: Read replicas for read-heavy hot spots
- Sharding by access pattern: Separate hot and cold data
3. Over-Complicating the Design
Start simple, add complexity only when justified:
"I'll start with a straightforward design and then discuss optimizations for scale..."
4. Forgetting to Justify Choices
Don't just draw boxes. Explain why:
"I'm using Kafka here instead of SQS because we need exactly-once semantics and the ability to replay events..."
If you can't explain why a component is there, remove it. Interviewers notice when candidates add components they don't understand.
Presenting Your Design
Structure your walkthrough:
- Start with the happy path: "Here's how a user posts a tweet..."
- Trace the data flow: "The request goes through the load balancer to..."
- Highlight key decisions: "I chose Redis for caching because..."
- Acknowledge trade-offs: "The downside of this approach is..."
- Invite questions: "Does this make sense? Should I dive deeper into any component?"
Draw as you talk. A silent candidate drawing is awkward. Narrate your thought process continuously.
What's Next
Your high-level design gives the blueprint. Now it's time for deep dives, where you'll be asked to zoom into specific components and solve detailed problems.