NEW: ML Mock & Coaching now available

Building Blocks

Message Queues

Understanding message queues, when to use them, and choosing between Kafka, RabbitMQ, and SQS.

6 min read

Message queues enable asynchronous communication between services. They decouple producers from consumers, improve reliability, and help handle traffic spikes.

Why Message Queues Matter

Without message queues:

  • Services are tightly coupled
  • Failures cascade between services
  • Traffic spikes can overwhelm downstream services
  • No way to retry failed operations

With message queues:

  • Services communicate without direct dependencies
  • Failed messages can be retried
  • Traffic is smoothed out over time
  • Better fault isolation
Info

If a user action doesn't need an immediate response, it's a candidate for async processing via a message queue.

Core Concepts

Producer and Consumer

Producer → Message Queue → Consumer

Example:
Order Service → Queue → Email Service
                     → Inventory Service
                     → Analytics Service

Message

The unit of data being transmitted. Contains:

  • Payload: The actual data (JSON, Protobuf, etc.)
  • Metadata: Timestamps, IDs, routing info
  • Headers: Additional context

Queue vs Topic

Queue vs Topic
NameDescription
Queue (Point-to-Point)Each message consumed by exactly one consumer. Good for task distribution.
Topic (Pub/Sub)Each message can be consumed by multiple subscribers. Good for event broadcasting.

Message Queue Technologies

Popular Message Queue Solutions
NameDescription
Apache KafkaDistributed streaming platform. High throughput, durable, replayable. Best for: event streaming, logs, high-volume data pipelines.
RabbitMQTraditional message broker. Rich routing, multiple protocols. Best for: complex routing, RPC patterns, lower volume.
Amazon SQSFully managed queue service. Simple, scalable, no ops. Best for: AWS-native apps, simple queuing needs.
Amazon SNSPub/sub notification service. Pairs with SQS. Best for: fan-out patterns, push notifications.
Redis StreamsRedis-based streaming. Good for existing Redis users. Best for: simpler streaming needs, lower latency.

Kafka vs RabbitMQ vs SQS

Message Queue Comparison
NameDescription
ThroughputKafka: Millions/sec. RabbitMQ: Tens of thousands/sec. SQS: Unlimited (managed).
Message retentionKafka: Configurable (days/weeks). RabbitMQ: Until consumed. SQS: Up to 14 days.
Replay capabilityKafka: Yes (core feature). RabbitMQ: No. SQS: No.
OrderingKafka: Per partition. RabbitMQ: Per queue. SQS: FIFO queues only.
OperationsKafka: Complex. RabbitMQ: Moderate. SQS: None (managed).

Common Patterns

Work Queue (Task Distribution)

Distribute tasks across multiple workers.

           ┌→ Worker 1
Producer → Queue → Worker 2
           └→ Worker 3

Use case: Image processing, email sending, report generation

Key considerations:

  • Ensure tasks are idempotent (safe to retry)
  • Use acknowledgments to prevent message loss
  • Implement dead-letter queues for failed tasks

Publish/Subscribe (Event Broadcasting)

Broadcast events to multiple interested services.

                    ┌→ Email Service
Order Service → Topic → Inventory Service
                    └→ Analytics Service

Use case: Event-driven architectures, notifications

Request/Reply (Async RPC)

Asynchronous request-response pattern.

Client → Request Queue → Server
      ← Reply Queue ←

Use case: Long-running operations where client needs the result

Saga Pattern (Distributed Transactions)

Coordinate transactions across multiple services.

Order Created → Payment Service (charge)
             → Inventory Service (reserve)
             → Shipping Service (schedule)

If any fails → Compensating transactions (refund, unreserve, cancel)
Challenge

Choose the Right Pattern

You're building an e-commerce system. When a user places an order, you need to: charge their card, update inventory, send confirmation email, and update analytics. How would you design this?

See recommended approach

Approach: Event-driven with Saga pattern

  1. Order Service validates order and publishes "OrderCreated" event
  2. Payment Service subscribes, charges card, publishes "PaymentCompleted" or "PaymentFailed"
  3. Inventory Service subscribes to "PaymentCompleted", reserves items
  4. Email Service subscribes to "PaymentCompleted", sends confirmation
  5. Analytics Service subscribes to all events for tracking

Why this works:

  • Services are decoupled
  • Each service can scale independently
  • Failures are isolated
  • Analytics gets complete event stream for analysis

Saga coordination:

  • If payment fails → no further processing needed
  • If inventory fails → trigger refund via compensating transaction
  • Use correlation ID to track all events for one order

Technology choice: Kafka for durability and replay capability, or SNS+SQS if on AWS and prefer managed services.

Delivery Guarantees

At-Most-Once

Message may be lost, never duplicated.

Producer sends → Network fails → Message lost

Use when: Loss is acceptable (metrics, logs that can be approximated)

At-Least-Once

Message delivered at least once, may be duplicated.

Producer sends → Consumer processes → Ack lost → Consumer reprocesses

Use when: Loss is unacceptable, consumers are idempotent

Exactly-Once

Message delivered exactly once. Hard to achieve.

Requires: Idempotent consumers OR transactional guarantees

Use when: Financial transactions, inventory updates

Warning

True exactly-once is complex and expensive. Most systems achieve "effectively exactly-once" through idempotent consumers with at-least-once delivery.

Handling Failures

Dead Letter Queues (DLQ)

Messages that fail processing go to a separate queue for investigation.

Main Queue → Consumer → Success
          ↘ Failure (after N retries) → Dead Letter Queue → Manual review

Retry Strategies

# Exponential backoff with jitter
def get_retry_delay(attempt):
    base_delay = 1  # seconds
    max_delay = 300  # 5 minutes
    jitter = random.uniform(0, 1)
    delay = min(base_delay * (2 ** attempt) + jitter, max_delay)
    return delay

Idempotency

Make consumers safe to run multiple times.

def process_order(order_id):
    # Check if already processed
    if db.exists(f"processed:{order_id}"):
        return  # Already done, skip

    # Process the order
    do_work(order_id)

    # Mark as processed
    db.set(f"processed:{order_id}", True)

Level-Based Expectations

Message Queue Knowledge by Level
NameDescription
Mid-Level (L4)Know when to use async processing. Understand basic queue concepts. Can explain producer/consumer pattern.
Senior (L5)Choose appropriate queue technology with reasoning. Discuss delivery guarantees. Design for failure (DLQ, retries, idempotency).
Staff+ (L6+)Design event-driven architectures. Handle exactly-once semantics. Discuss operational concerns (monitoring, backpressure, scaling consumers).
Engineering ManagerEvaluate build vs buy for messaging infrastructure. Plan team ownership of event contracts. Understand debugging complexity in async systems.

Common Interview Questions

"When would you use a message queue?"

"I'd use a message queue when the caller doesn't need an immediate response. For example: sending emails after signup, processing uploaded images, or updating analytics. This decouples services, handles failures gracefully, and smooths out traffic spikes."

"How do you ensure messages aren't lost?"

"Multiple layers: persistent storage in the queue (Kafka writes to disk), acknowledgments from consumers (only ack after successful processing), and dead letter queues for messages that fail repeatedly. Producers should retry on failure, and consumers should be idempotent."

"How do you handle message ordering?"

"It depends on the technology. In Kafka, ordering is guaranteed within a partition—so I'd partition by the entity that needs ordering (e.g., user_id for user events). In SQS, I'd use FIFO queues with message group IDs. For RabbitMQ, single-consumer queues maintain order."

What's Next

Message queues handle async communication. Next, we'll look at blob storage for handling large files and media content.