NEW: ML Mock & Coaching now available

Building Blocks

Load Balancers

Understanding load balancing strategies, L4 vs L7, and common algorithms for distributing traffic.

7 min read

Load balancers distribute incoming traffic across multiple servers, enabling horizontal scaling and improving reliability. They're present in virtually every system design.

Why Load Balancers Matter

Without load balancing:

  • Single server handles all traffic (single point of failure)
  • Can't scale horizontally
  • No graceful handling of server failures
  • Uneven resource utilization

With load balancing:

  • Traffic distributed across multiple servers
  • Failed servers automatically removed
  • Easy to add/remove capacity
  • Better resource utilization
Info

In interviews, you'll almost always have a load balancer. The key is knowing which type and algorithm to use.

Types of Load Balancers

L4 (Transport Layer) Load Balancers

Operate at the TCP/UDP level. Route based on IP address and port.

How it works:

  1. Client connects to load balancer IP
  2. Load balancer selects a backend server
  3. All packets in the connection go to the same server

Pros:

  • Very fast (no packet inspection)
  • Protocol agnostic
  • Lower resource usage

Cons:

  • Can't make decisions based on content
  • Limited health checking

Examples: AWS NLB, HAProxy (L4 mode), Linux IPVS

L7 (Application Layer) Load Balancers

Operate at the HTTP/HTTPS level. Can inspect and route based on content.

How it works:

  1. Client establishes connection to load balancer
  2. Load balancer reads HTTP request
  3. Routes based on URL, headers, cookies, etc.
  4. Maintains separate connection to backend

Pros:

  • Content-based routing (URL path, headers)
  • SSL termination
  • Request modification
  • Better health checks

Cons:

  • Higher latency (must parse request)
  • More resource intensive
  • Terminates SSL (security consideration)

Examples: AWS ALB, Nginx, HAProxy (L7 mode), Envoy

L4 vs L7 Load Balancers
NameDescription
SpeedL4: Faster (no parsing). L7: Slightly slower (parses HTTP).
Routing flexibilityL4: IP/port only. L7: URL, headers, cookies, query params.
SSL terminationL4: Passthrough. L7: Terminates and re-encrypts.
Health checksL4: TCP connect. L7: HTTP status codes, response content.
Use caseL4: Raw performance, non-HTTP. L7: Web applications, microservices.

Load Balancing Algorithms

Round Robin

Distribute requests sequentially across servers.

Request 1 → Server A
Request 2 → Server B
Request 3 → Server C
Request 4 → Server A
...

Pros: Simple, even distribution Cons: Doesn't account for server capacity or current load

Weighted Round Robin

Like round robin, but servers have different weights.

Server A (weight 3): Gets 3x traffic
Server B (weight 1): Gets 1x traffic

Requests: A, A, A, B, A, A, A, B, ...

Use when: Servers have different capacities.

Least Connections

Route to the server with fewest active connections.

Server A: 10 connections →
Server B: 5 connections  → New request goes here
Server C: 8 connections  →

Pros: Adapts to server load Cons: Doesn't account for connection "weight" (some requests are heavier)

Weighted Least Connections

Combines weights with connection counting.

Use when: Mixed server capacities with varying request complexity.

IP Hash

Hash client IP to consistently route to the same server.

hash(client_ip) % num_servers = target_server

Pros: Session affinity without cookies Cons: Uneven distribution if some IPs are busier

Least Response Time

Route to server with fastest response time.

Pros: Optimizes for latency Cons: Requires continuous monitoring, can oscillate

Challenge

Choose the Right Algorithm

You're designing a web application where users have sessions stored in server memory. Which load balancing algorithm would you use?

See analysis

Challenge: Session data is stored on individual servers. If a user's request goes to a different server, they lose their session.

Options:

  1. IP Hash: Routes same IP to same server. Works but breaks with NAT (multiple users behind same IP).

  2. Cookie-based sticky sessions (L7): Load balancer adds a cookie identifying the server. Most reliable for web applications.

  3. Better architecture: Externalize sessions to Redis. Then any algorithm works, and you get session persistence across server restarts.

Recommendation: Use sticky sessions as a short-term solution, but architect for externalized state. In the interview, mention both the quick fix and the better long-term solution.

Health Checks

Load balancers need to know which servers are healthy.

Passive Health Checks

Monitor actual traffic for errors.

If server returns 5xx errors → mark unhealthy
If connection fails → mark unhealthy

Pros: No extra traffic Cons: Only detects failures after they affect users

Active Health Checks

Periodically probe servers.

Every 10 seconds:
  GET /health → expect 200 OK
  If 3 consecutive failures → remove from pool
  If 2 consecutive successes → add back to pool

Pros: Detects issues before they affect users Cons: Extra traffic, need to implement health endpoint

Health Check Best Practices

# Good health check endpoint
@app.route('/health')
def health():
    # Check critical dependencies
    db_ok = check_database_connection()
    cache_ok = check_cache_connection()

    if db_ok and cache_ok:
        return {'status': 'healthy'}, 200
    else:
        return {'status': 'unhealthy', 'db': db_ok, 'cache': cache_ok}, 503

Load Balancer Patterns

Global Load Balancing (GSLB)

Distribute traffic across regions/data centers.

User in Europe → DNS → European data center
User in Asia → DNS → Asian data center

Technologies: AWS Route 53, Cloudflare, Akamai

Multi-Tier Load Balancing

Multiple layers of load balancers.

Internet → Global LB (DNS)
        → Regional LB (L4)
        → Service LB (L7)
        → Application servers

Service Mesh Load Balancing

Sidecar proxies handle service-to-service load balancing.

Service A → Envoy sidecar → Envoy sidecar → Service B

Technologies: Istio, Linkerd, Consul Connect

Level-Based Expectations

Load Balancing Knowledge by Level
NameDescription
Mid-Level (L4)Know to put a load balancer in front of servers. Understand round robin. Can explain why LBs improve reliability.
Senior (L5)Distinguish L4 vs L7 with trade-offs. Choose appropriate algorithm for the use case. Discuss health checks and failure handling.
Staff+ (L6+)Design multi-tier LB architectures. Discuss global load balancing and DNS. Handle edge cases like thundering herd on server recovery.
Engineering ManagerEvaluate LB costs and operational complexity. Plan capacity and scaling. Understand SLA implications of LB choices.

Common Interview Scenarios

"How do you handle a server failure?"

"The load balancer continuously runs health checks. When a server fails the health check—say, three consecutive failures—the LB removes it from the pool. Existing connections are drained or terminated. Traffic automatically routes to healthy servers. When the server recovers and passes health checks, it's added back to the pool."

"How do you scale to handle 10x traffic?"

"With a load balancer in place, I'd add more application servers behind it. The LB automatically distributes traffic to new servers. I'd also consider if the LB itself needs scaling—using multiple LBs with DNS round-robin or a global load balancer in front."

"How do you do zero-downtime deployments?"

"Using the load balancer's connection draining. During deployment: remove old server from LB pool, wait for existing connections to complete, deploy new version, add back to pool after health check passes. Repeat for each server."

Tip

In interviews, the load balancer is often drawn as a single box, but mention that in production it would be redundant (multiple LBs) to avoid being a single point of failure.

What's Next

Load balancers handle synchronous request distribution. Next, we'll look at message queues for asynchronous communication between services.