Batch Inference API
Request batching, GPU routing, and latency SLOs
About this session
A 90-minute live deep-dive into designing a batch inference API: aggregate synchronous HTTP requests into batches, route them to GPU workers, map responses back to callers, and hold a tight latency budget while maximizing throughput at 10,000 RPS over a fixed backend function. Great prep for AI-infra interviews at Anthropic and OpenAI.
Course Syllabus
Course Overview
- Format:
- Live Zoom, 90 minutes (75 min teaching + 15 min Q&A)
- Audience:
- Mid-level to Staff+ engineers prepping for AI-infra interviews (Anthropic, OpenAI)
- Prerequisite:
- Comfort with HTTP request/response, basic concurrency, and how GPUs batch work
What You Will Learn
- Aggregate synchronous HTTP requests into batches and map responses back to the original callers.
- Route batches to available GPU workers while holding a tight p99 latency budget.
- Derive the throughput vs. latency tradeoff for batch size and max-wait windows live.
- Handle the fixed `batchstring(inputs: list[str]) -> list[str]` backend at 10,000 RPS.
- Survive worker failure and backpressure without dropping in-flight requests.
What This Course Is NOT
- Not model training — this is inference serving only
- Not a memorization exercise — you'll derive batch-size math live
Pre-Class Preparation (24 hours before)
Read the problem statement only (10 min):
Design a service that accepts individual synchronous HTTP requests, aggregates them into batches internally, routes batches to available GPU workers, maps responses back to the original requesters, and keeps latency low while maximizing throughput at 10,000 RPS. The backend function `batchstring(inputs: list[str]) -> list[str]` is fixed and cannot be modified.
Think about (don't research yet):
- What's the simplest batching trigger — size, time, or both?
- Where does a request wait, and what's the worst-case added latency?
- What happens to a half-filled batch when a GPU worker dies?
Come to class with a first-pass batch size and max-wait number written down.
Lecture: author of showoffer
Jack
ShowOffer Coach
12+ years of experience in AI/ML infrastructure and distributed aggregation systems. Tech lead for fan-out architectures serving millions of product queries.
Reviews
Reviewer 006U
Jun 18, 2026
Pre-session Q&A
Questions confirmed attendees asked before this session.
No questions yet.