Load balancers distribute incoming traffic across multiple servers, enabling horizontal scaling and improving reliability. They're present in virtually every system design.
Why Load Balancers Matter
Without load balancing:
- Single server handles all traffic (single point of failure)
- Can't scale horizontally
- No graceful handling of server failures
- Uneven resource utilization
With load balancing:
- Traffic distributed across multiple servers
- Failed servers automatically removed
- Easy to add/remove capacity
- Better resource utilization
In interviews, you'll almost always have a load balancer. The key is knowing which type and algorithm to use.
Types of Load Balancers
L4 (Transport Layer) Load Balancers
Operate at the TCP/UDP level. Route based on IP address and port.
How it works:
- Client connects to load balancer IP
- Load balancer selects a backend server
- All packets in the connection go to the same server
Pros:
- Very fast (no packet inspection)
- Protocol agnostic
- Lower resource usage
Cons:
- Can't make decisions based on content
- Limited health checking
Examples: AWS NLB, HAProxy (L4 mode), Linux IPVS
L7 (Application Layer) Load Balancers
Operate at the HTTP/HTTPS level. Can inspect and route based on content.
How it works:
- Client establishes connection to load balancer
- Load balancer reads HTTP request
- Routes based on URL, headers, cookies, etc.
- Maintains separate connection to backend
Pros:
- Content-based routing (URL path, headers)
- SSL termination
- Request modification
- Better health checks
Cons:
- Higher latency (must parse request)
- More resource intensive
- Terminates SSL (security consideration)
Examples: AWS ALB, Nginx, HAProxy (L7 mode), Envoy
| L4 vs L7 Load Balancers | |
|---|---|
| Name | Description |
Speed | L4: Faster (no parsing). L7: Slightly slower (parses HTTP). |
Routing flexibility | L4: IP/port only. L7: URL, headers, cookies, query params. |
SSL termination | L4: Passthrough. L7: Terminates and re-encrypts. |
Health checks | L4: TCP connect. L7: HTTP status codes, response content. |
Use case | L4: Raw performance, non-HTTP. L7: Web applications, microservices. |
Load Balancing Algorithms
Round Robin
Distribute requests sequentially across servers.
Request 1 → Server A
Request 2 → Server B
Request 3 → Server C
Request 4 → Server A
...
Pros: Simple, even distribution Cons: Doesn't account for server capacity or current load
Weighted Round Robin
Like round robin, but servers have different weights.
Server A (weight 3): Gets 3x traffic
Server B (weight 1): Gets 1x traffic
Requests: A, A, A, B, A, A, A, B, ...
Use when: Servers have different capacities.
Least Connections
Route to the server with fewest active connections.
Server A: 10 connections →
Server B: 5 connections → New request goes here
Server C: 8 connections →
Pros: Adapts to server load Cons: Doesn't account for connection "weight" (some requests are heavier)
Weighted Least Connections
Combines weights with connection counting.
Use when: Mixed server capacities with varying request complexity.
IP Hash
Hash client IP to consistently route to the same server.
hash(client_ip) % num_servers = target_server
Pros: Session affinity without cookies Cons: Uneven distribution if some IPs are busier
Least Response Time
Route to server with fastest response time.
Pros: Optimizes for latency Cons: Requires continuous monitoring, can oscillate
Choose the Right Algorithm
You're designing a web application where users have sessions stored in server memory. Which load balancing algorithm would you use?
See analysis
Challenge: Session data is stored on individual servers. If a user's request goes to a different server, they lose their session.
Options:
-
IP Hash: Routes same IP to same server. Works but breaks with NAT (multiple users behind same IP).
-
Cookie-based sticky sessions (L7): Load balancer adds a cookie identifying the server. Most reliable for web applications.
-
Better architecture: Externalize sessions to Redis. Then any algorithm works, and you get session persistence across server restarts.
Recommendation: Use sticky sessions as a short-term solution, but architect for externalized state. In the interview, mention both the quick fix and the better long-term solution.
Health Checks
Load balancers need to know which servers are healthy.
Passive Health Checks
Monitor actual traffic for errors.
If server returns 5xx errors → mark unhealthy
If connection fails → mark unhealthy
Pros: No extra traffic Cons: Only detects failures after they affect users
Active Health Checks
Periodically probe servers.
Every 10 seconds:
GET /health → expect 200 OK
If 3 consecutive failures → remove from pool
If 2 consecutive successes → add back to pool
Pros: Detects issues before they affect users Cons: Extra traffic, need to implement health endpoint
Health Check Best Practices
# Good health check endpoint
@app.route('/health')
def health():
# Check critical dependencies
db_ok = check_database_connection()
cache_ok = check_cache_connection()
if db_ok and cache_ok:
return {'status': 'healthy'}, 200
else:
return {'status': 'unhealthy', 'db': db_ok, 'cache': cache_ok}, 503
Load Balancer Patterns
Global Load Balancing (GSLB)
Distribute traffic across regions/data centers.
User in Europe → DNS → European data center
User in Asia → DNS → Asian data center
Technologies: AWS Route 53, Cloudflare, Akamai
Multi-Tier Load Balancing
Multiple layers of load balancers.
Internet → Global LB (DNS)
→ Regional LB (L4)
→ Service LB (L7)
→ Application servers
Service Mesh Load Balancing
Sidecar proxies handle service-to-service load balancing.
Service A → Envoy sidecar → Envoy sidecar → Service B
Technologies: Istio, Linkerd, Consul Connect
Level-Based Expectations
| Load Balancing Knowledge by Level | |
|---|---|
| Name | Description |
Mid-Level (L4) | Know to put a load balancer in front of servers. Understand round robin. Can explain why LBs improve reliability. |
Senior (L5) | Distinguish L4 vs L7 with trade-offs. Choose appropriate algorithm for the use case. Discuss health checks and failure handling. |
Staff+ (L6+) | Design multi-tier LB architectures. Discuss global load balancing and DNS. Handle edge cases like thundering herd on server recovery. |
Engineering Manager | Evaluate LB costs and operational complexity. Plan capacity and scaling. Understand SLA implications of LB choices. |
Common Interview Scenarios
"How do you handle a server failure?"
"The load balancer continuously runs health checks. When a server fails the health check—say, three consecutive failures—the LB removes it from the pool. Existing connections are drained or terminated. Traffic automatically routes to healthy servers. When the server recovers and passes health checks, it's added back to the pool."
"How do you scale to handle 10x traffic?"
"With a load balancer in place, I'd add more application servers behind it. The LB automatically distributes traffic to new servers. I'd also consider if the LB itself needs scaling—using multiple LBs with DNS round-robin or a global load balancer in front."
"How do you do zero-downtime deployments?"
"Using the load balancer's connection draining. During deployment: remove old server from LB pool, wait for existing connections to complete, deploy new version, add back to pool after health check passes. Repeat for each server."
In interviews, the load balancer is often drawn as a single box, but mention that in production it would be redundant (multiple LBs) to avoid being a single point of failure.
What's Next
Load balancers handle synchronous request distribution. Next, we'll look at message queues for asynchronous communication between services.