NEW: ML Mock & Coaching now available

Questions

Distributed Counters

GoogleMetaNetflix

Design YouTube View Counter - This system design covers distributed counting at scale, idempotent event ingestion, hotspot mitigation, and near real-time serving with strong correctness guarantees.

40 min read

Staff+ Engineer: 10+ years of experience in low-latency streaming and AI infrastructure. Tech lead for real-time LLM serving platforms handling millions of concurrent requests. Engineering Manager: Over 12 years of industry experience with 6 years in engineering leadership, building and scaling tech infrastructure teams that deliver end-to-end large-scale distributed systems.

Challenge

Challenge Yourself

If you're still reading, you're already ahead — here are the real questions that turn "a counter" into a production-grade distributed system:

  1. Concurrency & Race Conditions

    How do we guarantee correct counter updates when many servers increment the same video counter at the same time, without relying on a single bottleneck or global lock?

  2. Dedup & View Semantics

    Given that a "view" has a real definition (watch-time threshold, retries, refreshes), how do we design a dedup + idempotency model so each legitimate view is counted exactly once?

  3. Near Real-time + Accuracy

    How do we deliver a near-real-time view count on the watch page while still ensuring the number is auditable and correct under late events, backfills, and fraud corrections?

Problem Statement

Design a backend system that powers YouTube long-form video view counters.

Key user behaviors (what the system should provide)

  1. Watch a video → the system records viewing activity and eventually increments the video's view count.
  2. Refresh / reconnect / retry (mobile drops, browser refresh, client retries) → the system avoids double-counting and still converges to the right number.
  3. Watch a popular (viral) video → the view count remains available and updates smoothly even under massive traffic spikes.
Out of Scope
  • Shorts view counting rules, ads view measurement, monetization/billing counters
  • Creator analytics beyond the public counter (unique viewers, watch time, retention curves)
  • Recommendation/search ranking and user personalization

Functional Requirements

FR1 — Ingest view events

Accept view signals for a video_id and persist them so the system can later decide whether a view should be counted.

FR2 — Serve view counts (watch page)

Provide GetViewCount(video_id) to return the public view count quickly and reliably for end users.

Non-Functional Requirements

NFR1 — Correctness

  • No double-counting on retries: the same logical view must be counted at most once (target: 0 duplicate increments per idempotency/dedup key).
  • No lost updates under concurrency: when many writes hit the same video_id (including the same shard), increments must not overwrite each other.

NFR2 — Latency / Freshness

  • GetViewCount(video_id) latency: p95 ≤ 50 ms, p99 ≤ 150 ms
  • View count freshness (staleness of displayed number): p95 ≤ 2 s, p99 ≤ 5 s

NFR3 — Scale / Hotspot Resistance

  • Peak ingest capacity: ≥ 200K view-events/sec (example target)
  • Hot video tolerance: ≥ 50K increments/sec for a single video_id without a single partition/shard melting
  • Read scale: support very high QPS for GetViewCount via caching/edge serving

NFR4 — Availability > Consistency

  • For view counts we choose AP: always accept view events and always serve a count, even during partitions, and we rely on event-log + rollups to converge to the correct number.
Challenge Yourself — Why AP is the Right CAP Choice Here?
  • Partition tolerance is non-negotiable. In a global system, cross-region partitions and partial network failures will happen.
  • When a partition happens, you don't want the watch page to break. Returning some view count is far better than timing out or erroring.
  • For views, the product can tolerate temporary staleness. Users don't require the number to be globally exact to the second.

So under partition:

  • Reads stay available (serve last known count).
  • Writes stay available (continue accepting view events regionally).
  • Counts converge later once replication/rollup catches up.

What we "give up" (and how we contain it)?

  • We give up strong consistency (no global linearizable counter).
  • We contain the impact by:
    • using an event log as the source of truth,
    • doing regional ingestion + aggregation,
    • and rolling up behind a watermark so the number converges predictably.

Requirement Summary

Functional Requirements (FRs)
NameDescription
1. Ingest View EventsAccept view signals for a video_id and persist them so the system can later decide whether a view should be counted.
2. Serve View CountsProvide GetViewCount(video_id) to return the public view count quickly and reliably for end users.
Non-Functional Requirements (NFRs)
NameDescription
1. CorrectnessNo double-counting on retries; no lost updates under concurrency.
2. Latency / FreshnessGetViewCount p95 ≤ 50 ms, p99 ≤ 150 ms; freshness p95 ≤ 2 s, p99 ≤ 5 s.
3. Scale / Hotspot ResistancePeak ingest ≥ 200K events/sec; hot video ≥ 50K increments/sec for a single video_id.
4. Availability > ConsistencyAP: always accept events and serve counts, even during partitions; converge via event-log + rollups.

Sign in to continue reading

"Distributed Counters" requires a free account to access.

Sign in to continue