Think Beyond the Happy Path
Leaderboard systems aren't just about sorting scores — they're about handling billions of updates per day, avoiding hot shards, and serving personalized rankings (global + friends) with sub-100ms latency.
Before diving into the material, take a moment to ask yourself:
- Do you know how to maintain a global leaderboard at billion updates/day without creating a hot shard that becomes a bottleneck?
- Do you know how to fetch K-neighbors around a user (players ranked just above and below them) efficiently for both global and friends leaderboards?
- Do you know how to handle hotspots at the top — what happens when millions of users query the same top-10 players simultaneously?
- Do you know how to recover from cache or ranking loss — can you rebuild the leaderboard quickly without re-processing all historical data?
You don't need to answer all of these right away. A Senior/Staff+ Engineer doesn't stop at basic functionality — they anticipate edge cases, design for resilience, and push for production-grade reliability.
But great systems start with great questions. What would you ask next?
What to Design?
Design a real-time leaderboard service for an online game that:
- Continuously ingests players' scores.
- Lets players see how they rank:
- Globally (top-N and current rank).
- Locally around them (K players above and below).
- Within their social graph (friends-only leaderboard).
- Within time windows such as daily, weekly, and seasonal leaderboards.
What will be your key assumptions during your interview and align with your interviewers?
Here are some examples we made, and the following design align these assumptions:
1. Game & Score Model
- Score is an integer.
- Each player has one active highest score (i.e., mode = "HIGH_SCORE") per leaderboard (per game × season × region × period type) ⇒ Score is non-decreasing within a leaderboard period
- System is multi-game and multi-leaderboard (e.g., different modes), but each leaderboard is keyed separately.
2. Scale & Traffic Model
- 100M DAU.
- Writes:
- 100–300M score updates / day.
- Peak: ~10K score updates / second (across all regions).
- Reads:
- ~5× write volume → ≈ 1B leaderboard reads / day.
- Mix of:
- "My rank + K neighbors".
- Top-N (global / regional / time-windowed).
- Friends-only leaderboards.
3. Leaderboard Scope (Multi-Dimension Support)
-
Global leaderboard
Per game / per season, across all regions.
-
Regional leaderboard
Regional splits such as NA / EU / APAC.
-
Time-window leaderboard
- Daily / weekly / monthly leaderboards.
- Each period starts at the exact first day (or boundary) of the period (no rolling windows).
-
Friends leaderboard
Rankings restricted to the user's friends list (social graph provided by a separate service).
Functional Requirements
FR1 – Ingest & Update Player Scores
- The system must accept score updates for a player on a given leaderboard.
- Input includes:
player_id,leaderboard_id(game / mode / region / period), andnew_scoreplus an idempotency key (e.g.,match_id). - The system must maintain a non-decreasing integer score per
(player_id, leaderboard_id).
FR2 – View My Global Rank & K Neighbors
- Given
player_idandleaderboard_id, the system must return:- The player's current score.
- The player's current global rank.
- Up to K players above and K players below the user on the global leaderboard.
- Must handle edge cases (player near top/bottom → fewer than K neighbors).
- Results should reflect recent updates in near real-time (tied to our <500 ms update SLO).
FR3 – View K Friends Around Me
- Given
player_id,leaderboard_id, andK, the system must return:- The player's rank and score within their friends-only leaderboard.
- Up to K friends above and K friends below them.
- Uses an external friends/social graph as input; our service filters or indexes rankings based on that list.
FR4 – View Top-N Leaderboards (Global / Regional / Time-Windowed)
- Given a
leaderboard_idandN, the system must return the top-N entries (player_id, score, rank) for that leaderboard. - Must support multiple scopes via
leaderboard_id:- Global (per game / per season, all regions).
- Regional (e.g., NA / EU / APAC).
- Time-windowed (daily / weekly / monthly / seasonal, each with a strict start date).
- Used for classic "Top 100 / Top 1000" views and tournament pages, and should refresh near real-time as scores change.
Why these 4 Functional Requirements matter?
These aren't just "screens" the product wants — each FR encodes a distinct access pattern that drives the whole design:
- FR1 (Ingest score) → ultra-high write throughput and idempotent updates per
(player_id, leaderboard_id). - FR2 (My global rank + K neighbors) → fast range queries around a single user on a sorted index (the core "ranked set" problem).
- FR3 (My rank among friends) → intersection of the global ranking with a dynamic friend set (social-graph flavored ranking).
- FR4 (Top-N by scope/time window) → hot prefix reads on a sorted structure, plus multi-dimensional keys (game × region × period).
If you get these four access patterns right, most follow-up questions (sharding, caching, K-neighbors, hotspots, recovery) become specific implementation choices, not surprises.
Non-Functional Requirements
NFR1 – Correctness & Determinism
- A player's score on a leaderboard is never lost, mis-ordered, or double-counted.
- Updates are monotonic and idempotent per
(player_id, leaderboard_id)(e.g., keyed bymatch_id).
NFR2 – Real-time Freshness & Latency
- Write → visible rank (unchanged):
- p95 ≤ 500 ms, p99 ≈ 1 s for the player's own score.
- Read APIs (my rank, neighbors, top-N), server-side:
- p95 ≤ 100 ms, p99 ≤ 500 ms.
- It's fine if some pages (e.g., big top-N lists) sit closer to the tail; players experience this as "instant enough" as long as it's clearly under ~0.5–1 s end-to-end including network.
NFR3 – Availability & Graceful Degradation
- Leaderboard reads and score updates: ≥ 99.9% availability.
- Under failures, prefer serving last-known data (cached rank / top-N) over errors.
- If degraded, surface simple hints like "last updated X seconds ago" instead of breaking UX.
NFR4 – Scalability & Hotspot Resistance
- Handle 100–300M score updates/day → peak around 10K updates/sec.
- Handle ~1B leaderboard reads/day (~5× writes).
- Sharding / caching must avoid any single hot shard (e.g., top of global board or celebrity profiles) becoming a bottleneck.
- System scales horizontally by adding nodes, not by vertically beefing up a single box.
Why This Order?
- NFR1 (Correctness) – If the leaderboard shows the wrong score or rank for you, the product loses trust immediately; correctness for your own score is non-negotiable.
- NFR2 (Freshness & Latency) – Once it's correct, it needs to feel real-time and snappy; otherwise players assume it's buggy or laggy.
- NFR3 (Availability) – A correct, fast leaderboard that's often down is still useless, so we keep it up and degrading gracefully.
- NFR4 (Scalability) – Finally, we make sure the same guarantees hold at 100M DAU and beyond, without hot shards melting the system.
High Level Design
What should be the best Delivery Framework during your interview?
For each important FR, we suggest structure the High-Level Design in this way:
- API (Abstract)
- Define who calls what and what the payload looks like.
- Focus on intent, not syntax: path, key fields, and behavior guarantees (idempotency, monotonicity, etc.).
- Why it's good in interviews: anchors the discussion in something concrete the product/client would actually use and proves I'm designing from the outside-in.
- Core Entities / Data Model
- List the minimal tables / objects needed to fulfill this FR:
- Primary keys.
- Key columns that encode correctness (e.g.,
score,last_match_id,last_update_ts).
- Don't over-model—only what this FR really needs.
- Why it's good: shows I can translate APIs into a clean data model and reason about invariants like non-decreasing scores and idempotency.
- List the minimal tables / objects needed to fulfill this FR:
- Workflow / Request Flow
- Walk through the end-to-end path for this FR on a simple system:
- Where the request enters (API service).
- What validations happen.
- Which entities are read/updated.
- What the response looks like.
- Why it's good: proves I can actually wire components together into a working system, not just list buzzword boxes.
- Walk through the end-to-end path for this FR on a simple system:
- Simple Diagram (Text or Boxes & Arrows)
- Draw a small diagram for this vertical slice:
- Client → API Service → Storage / Cache → back to Client.
- Keep it single-region and non-sharded at first; scale comes later in deep dives.
- Why it's good: gives the interviewer a visual anchor and makes follow-up questions ("where would you shard?", "where do you cache?") much easier to explore.
- Draw a small diagram for this vertical slice:
FR1 – Ingest & Update Player Scores
API (Abstract)
We'll assume the game backend calls us after each match.
- Endpoint (abstract)
POST /leaderboards/{leaderboard_id}/players/{player_id}/score
Content-Type: application/json
{
"new_score": 12345,
"match_id": "match-uuid-123", // idempotency key
}
- Response (example)
{
"player_id": "player-123",
"leaderboard_id": "lb-ranked-na-s12-week-2025-11-17",
"score": 12345,
"previous_score": 12000,
"last_match_id": "match-uuid-123", // last match that actually change sco
"updated_at": "2025-11-17T10:15:00Z",
"update_applied": true // false if this was a pure idempotent retry
}
Behavior guarantees
- Non-decreasing: if
new_score < current_score, we ignore the update or return a 4xx error (pick one and be explicit). - Idempotent: if we've already processed this
match_idfor this(player_id, leaderboard_id), the request is a no-op: we return the existing score andupdate_applied=false.
Core Entities
leaderboards
- Describes one concrete leaderboard: game + mode + region + season + time window.
- Tells us where a score update should go and whether that leaderboard is active.
- Encodes score policy so we know we're in the
HIGH_SCOREworld (monotonic best-score).
| leaderboards | |
|---|---|
| Name | Description |
leaderboard_id | Primary key; unique identifier for this leaderboard (game + mode + region + period). |
game_id | Which game this leaderboard belongs to. |
mode | Game mode type, e.g. `HIGH_SCORE` |
region | Region for this leaderboard: `GLOBAL`, `NA`, `EU`, `APAC`, etc. |
season_id | Season identifier (e.g., "S12"). |
period_type | Time window type: `DAILY`, `WEEKLY`, `MONTHLY`, `SEASONAL`. |
period_start_ts | Start timestamp of this leaderboard period. |
period_end_ts | End timestamp of this period (null while active). |
status | `ACTIVE`, `FROZEN`, or `ARCHIVED` (controls whether we accept writes). |
created_at | When this leaderboard was created. |
2. score_events (append-only log of raw updates)
- Logs every match event that tried to update a leaderboard.
- Useful for audit, debugging, and later replay / rebuild if needed.
- Enforces uniqueness on
(leaderboard_id, player_id, match_id)so the same match can't be ingested twice.
| score_events | |
|---|---|
| Name | Description |
event_id | Primary key; unique id for this score event. |
leaderboard_id | Which leaderboard this event targets (FK to `leaderboards`). |
player_id | Player whose score this event is about. |
match_id | Match identifier; used for idempotency (one event per `(leaderboard, player, match)`). |
new_score | Candidate high score from this match (before applying `HIGH_SCORE` logic). |
event_ts | When this event was recorded. |
3. leaderboard_scores (canonical per-player state)
- Holds the current visible score per
(player, leaderboard). - This is the row all future FRs will query for:
- My score, my rank, K neighbors, top-N, etc.
- Encodes the HIGH_SCORE, monotonic semantics and idempotency.
| leaderboard_scores | |
|---|---|
| Name | Description |
leaderboard_id | Leaderboard this score belongs to (PK part 1; FK to `leaderboards`). |
player_id | Player identifier (PK part 2). |
score | Current high score for this player on this leaderboard (monotonic under `HIGH_SCORE` policy). |
last_match_id | The last match that actually changed this high score (used for idempotency and debugging). |
last_update_ts | When this high score was last updated. |
version | Optional version counter for optimistic locking / concurrency control. |
Workflow
Step 1 - Match service sends score update
- After a match, the match service
POSTtheleaderboard_id,player_id,match_id,new_scoreto API-GW.
Step 2 - API-GW routes to Score Ingestion Service
- API-GW does auth / basic validation and forwards the request to the Score Ingestion Service.
Step 3 - Score Ingestion validates leaderboard & logs event
- Checks that the
leaderboard_idexists and isACTIVE. - Inserts a row into
score_events(leaderboard_id,player_id,match_id,new_score,event_ts). - The unique index on
(leaderboard_id, player_id, match_id)enforces idempotency (duplicate match → treated as retry).
Step 4 - Apply HIGH_SCORE rule using leaderboard_scores
- Reads current row for
(leaderboard_id, player_id)fromleaderboard_scores. - If
match_id == last_match_id→ retry, no change. - Else if
new_score <= current score→ not a new high score, no change. - Else → update
leaderboard_scoreswithscore = new_score,last_match_id = match_id,last_update_ts = now().
Step 5 - Respond back to match service
- Score Ingestion Service returns final
score,previous_score,last_match_id, andupdate_applied = true/false→ API-GW → match service.
Design Diagram

Click to expand