Message queues enable asynchronous communication between services. They decouple producers from consumers, improve reliability, and help handle traffic spikes.
Why Message Queues Matter
Without message queues:
- Services are tightly coupled
- Failures cascade between services
- Traffic spikes can overwhelm downstream services
- No way to retry failed operations
With message queues:
- Services communicate without direct dependencies
- Failed messages can be retried
- Traffic is smoothed out over time
- Better fault isolation
If a user action doesn't need an immediate response, it's a candidate for async processing via a message queue.
Core Concepts
Producer and Consumer
Producer → Message Queue → Consumer
Example:
Order Service → Queue → Email Service
→ Inventory Service
→ Analytics Service
Message
The unit of data being transmitted. Contains:
- Payload: The actual data (JSON, Protobuf, etc.)
- Metadata: Timestamps, IDs, routing info
- Headers: Additional context
Queue vs Topic
| Queue vs Topic | |
|---|---|
| Name | Description |
Queue (Point-to-Point) | Each message consumed by exactly one consumer. Good for task distribution. |
Topic (Pub/Sub) | Each message can be consumed by multiple subscribers. Good for event broadcasting. |
Message Queue Technologies
| Popular Message Queue Solutions | |
|---|---|
| Name | Description |
Apache Kafka | Distributed streaming platform. High throughput, durable, replayable. Best for: event streaming, logs, high-volume data pipelines. |
RabbitMQ | Traditional message broker. Rich routing, multiple protocols. Best for: complex routing, RPC patterns, lower volume. |
Amazon SQS | Fully managed queue service. Simple, scalable, no ops. Best for: AWS-native apps, simple queuing needs. |
Amazon SNS | Pub/sub notification service. Pairs with SQS. Best for: fan-out patterns, push notifications. |
Redis Streams | Redis-based streaming. Good for existing Redis users. Best for: simpler streaming needs, lower latency. |
Kafka vs RabbitMQ vs SQS
| Message Queue Comparison | |
|---|---|
| Name | Description |
Throughput | Kafka: Millions/sec. RabbitMQ: Tens of thousands/sec. SQS: Unlimited (managed). |
Message retention | Kafka: Configurable (days/weeks). RabbitMQ: Until consumed. SQS: Up to 14 days. |
Replay capability | Kafka: Yes (core feature). RabbitMQ: No. SQS: No. |
Ordering | Kafka: Per partition. RabbitMQ: Per queue. SQS: FIFO queues only. |
Operations | Kafka: Complex. RabbitMQ: Moderate. SQS: None (managed). |
Common Patterns
Work Queue (Task Distribution)
Distribute tasks across multiple workers.
┌→ Worker 1
Producer → Queue → Worker 2
└→ Worker 3
Use case: Image processing, email sending, report generation
Key considerations:
- Ensure tasks are idempotent (safe to retry)
- Use acknowledgments to prevent message loss
- Implement dead-letter queues for failed tasks
Publish/Subscribe (Event Broadcasting)
Broadcast events to multiple interested services.
┌→ Email Service
Order Service → Topic → Inventory Service
└→ Analytics Service
Use case: Event-driven architectures, notifications
Request/Reply (Async RPC)
Asynchronous request-response pattern.
Client → Request Queue → Server
← Reply Queue ←
Use case: Long-running operations where client needs the result
Saga Pattern (Distributed Transactions)
Coordinate transactions across multiple services.
Order Created → Payment Service (charge)
→ Inventory Service (reserve)
→ Shipping Service (schedule)
If any fails → Compensating transactions (refund, unreserve, cancel)
Choose the Right Pattern
You're building an e-commerce system. When a user places an order, you need to: charge their card, update inventory, send confirmation email, and update analytics. How would you design this?
See recommended approach
Approach: Event-driven with Saga pattern
- Order Service validates order and publishes "OrderCreated" event
- Payment Service subscribes, charges card, publishes "PaymentCompleted" or "PaymentFailed"
- Inventory Service subscribes to "PaymentCompleted", reserves items
- Email Service subscribes to "PaymentCompleted", sends confirmation
- Analytics Service subscribes to all events for tracking
Why this works:
- Services are decoupled
- Each service can scale independently
- Failures are isolated
- Analytics gets complete event stream for analysis
Saga coordination:
- If payment fails → no further processing needed
- If inventory fails → trigger refund via compensating transaction
- Use correlation ID to track all events for one order
Technology choice: Kafka for durability and replay capability, or SNS+SQS if on AWS and prefer managed services.
Delivery Guarantees
At-Most-Once
Message may be lost, never duplicated.
Producer sends → Network fails → Message lost
Use when: Loss is acceptable (metrics, logs that can be approximated)
At-Least-Once
Message delivered at least once, may be duplicated.
Producer sends → Consumer processes → Ack lost → Consumer reprocesses
Use when: Loss is unacceptable, consumers are idempotent
Exactly-Once
Message delivered exactly once. Hard to achieve.
Requires: Idempotent consumers OR transactional guarantees
Use when: Financial transactions, inventory updates
True exactly-once is complex and expensive. Most systems achieve "effectively exactly-once" through idempotent consumers with at-least-once delivery.
Handling Failures
Dead Letter Queues (DLQ)
Messages that fail processing go to a separate queue for investigation.
Main Queue → Consumer → Success
↘ Failure (after N retries) → Dead Letter Queue → Manual review
Retry Strategies
# Exponential backoff with jitter
def get_retry_delay(attempt):
base_delay = 1 # seconds
max_delay = 300 # 5 minutes
jitter = random.uniform(0, 1)
delay = min(base_delay * (2 ** attempt) + jitter, max_delay)
return delay
Idempotency
Make consumers safe to run multiple times.
def process_order(order_id):
# Check if already processed
if db.exists(f"processed:{order_id}"):
return # Already done, skip
# Process the order
do_work(order_id)
# Mark as processed
db.set(f"processed:{order_id}", True)
Level-Based Expectations
| Message Queue Knowledge by Level | |
|---|---|
| Name | Description |
Mid-Level (L4) | Know when to use async processing. Understand basic queue concepts. Can explain producer/consumer pattern. |
Senior (L5) | Choose appropriate queue technology with reasoning. Discuss delivery guarantees. Design for failure (DLQ, retries, idempotency). |
Staff+ (L6+) | Design event-driven architectures. Handle exactly-once semantics. Discuss operational concerns (monitoring, backpressure, scaling consumers). |
Engineering Manager | Evaluate build vs buy for messaging infrastructure. Plan team ownership of event contracts. Understand debugging complexity in async systems. |
Common Interview Questions
"When would you use a message queue?"
"I'd use a message queue when the caller doesn't need an immediate response. For example: sending emails after signup, processing uploaded images, or updating analytics. This decouples services, handles failures gracefully, and smooths out traffic spikes."
"How do you ensure messages aren't lost?"
"Multiple layers: persistent storage in the queue (Kafka writes to disk), acknowledgments from consumers (only ack after successful processing), and dead letter queues for messages that fail repeatedly. Producers should retry on failure, and consumers should be idempotent."
"How do you handle message ordering?"
"It depends on the technology. In Kafka, ordering is guaranteed within a partition—so I'd partition by the entity that needs ordering (e.g., user_id for user events). In SQS, I'd use FIFO queues with message group IDs. For RabbitMQ, single-consumer queues maintain order."
What's Next
Message queues handle async communication. Next, we'll look at blob storage for handling large files and media content.