NEW: ML Mock & Coaching now available

Questions

Webhook Delivery System

StripeOpenAI

Design a reliable webhook delivery system - handling event ingestion, guaranteed delivery, and fault tolerance at scale.

40 min read

Challenge

Think Beyond the Happy Path

Before diving into the solution, consider these critical questions:

  • How would you guarantee delivery when the receiving endpoint might be temporarily down?
  • How would you handle a "hot client" that suddenly receives 10x their normal event volume?
  • How would you prevent event tampering so clients can trust the payloads they receive?
  • How would you scale to billions of events per day while maintaining sub-second delivery latency?

You don't need to answer all of these right away. A Senior/Staff+ Engineer doesn't stop at basic functionality — they anticipate edge cases, design for resilience, and push for production-grade reliability.

But great systems start with great questions. What would you ask next?

Problem Statement

Design a webhook delivery system that enables real-time communication between applications by reliably delivering event notifications to registered endpoints.

A webhook is a mechanism that allows one application to notify another when an event occurs. Instead of constantly polling for updates, the receiving application simply waits for notifications to be pushed to its designated endpoint. This pattern is fundamental to modern integrations - from payment processing to CI/CD pipelines.

Real-World Example: Payment Processing

Consider an e-commerce platform like Amazon processing payments through Stripe:

  1. Customer initiates payment on Amazon
  2. Amazon sends payment request to Stripe
  3. Stripe processes payment with card networks (Visa, Mastercard)
  4. When payment succeeds or fails, Stripe sends a webhook to Amazon's callback URL
  5. Amazon receives the notification and updates the order status

Without webhooks, Amazon would need to continuously poll Stripe asking "Is the payment done yet?" - wasteful and slow. Webhooks flip this model: Stripe proactively notifies Amazon the moment something happens.

Clarifying Questions

Before designing, clarify these requirements with your interviewer:

Info

Key Questions to Ask:

  1. What types of events trigger webhooks? (user actions, system events, third-party integrations)
  2. Expected volume of events? (millions vs billions per day)
  3. Delivery guarantees needed? (at-least-once, exactly-once)
  4. Event ordering requirements? (strict ordering within a client, or best-effort)
  5. Latency requirements? (sub-second, minutes acceptable)
  6. Security requirements? (encryption in transit, payload signing)

The answers shape fundamental design decisions:

  • High throughput + low latency → Consider streaming (Kafka) over queuing (SQS)
  • Strong security requirements → Plan for HMAC signing and payload encryption
  • Strict ordering → Partition by client to maintain order within each client's events

Functional Requirements

FR1 - Register, update, and delete webhooks

Clients can register webhook endpoints with a callback URL, specifying which event types they want to receive. They can update configurations (URL, event filters, retry policies) or delete webhooks when no longer needed.

FR2 - Receive events through webhooks

When a source application generates an event (e.g., payment processed), the webhook system delivers it to all matching registered endpoints. Clients receive HTTP POST requests containing event details.

FR3 - Filter events by type

Clients can subscribe to specific event types (e.g., only payment.succeeded and payment.failed, ignoring payment.created). This reduces noise and processing overhead for clients.

Non-Functional Requirements

NFR1 - High Scalability

Support up to 1 billion events per day. This translates to:

  • Average: ~11,500 events per second
  • Peak (3x): ~35,000 events per second

NFR2 - High Availability

The system must remain available for receiving, processing, and delivering events even during partial failures. Source applications depend on reliable event ingestion.

NFR3 - Fault Tolerance with Retries

When delivery fails (client endpoint down, network issues), automatically retry with exponential backoff. After exhausting retries, move to dead letter queue for manual intervention.

NFR4 - Security

Prevent unauthorized access and data tampering. Events must be signed so clients can verify authenticity and integrity.

NFR5 - Durability

No event loss. Once an event is accepted, it must eventually be delivered (or explicitly failed after retry exhaustion).

Staff+ Signal: Quantifying Requirements

Great candidates don't just say "high scalability" - they quantify it:

  • 1B events/day → ~11.5K events/second average, ~35K peak
  • Delivery SLA: P95 under 5 seconds for online endpoints
  • Retry policy: 5 attempts over 24 hours with exponential backoff
  • Retention: Failed events kept 7 days for replay

Adding specific numbers shows you understand operational constraints and can make informed tradeoffs.

API Design

Register a Webhook

POST /api/v1/webhooks
Authorization: Bearer {JWT_TOKEN}

Request:
{
  "callback_url": "https://api.example.com/webhooks/payments",
  "event_types": ["payment.succeeded", "payment.failed"],
  "retry_config": {
    "max_retries": 5,
    "initial_delay_ms": 1000,
    "max_delay_ms": 30000,
    "backoff_multiplier": 2.0
  }
}

Response: 200 OK
{
  "webhook_id": "wh_abc123",
  "signing_secret": "whsec_xyz789",  // Client uses this to verify payloads
  "status": "active",
  "created_at": "2025-01-15T10:30:00Z"
}

The signing_secret is critical - clients use it to verify that incoming webhook payloads actually came from our system and weren't tampered with.

Sign in to continue reading

"Webhook Delivery System" requires a free account to access.

Sign in to continue