AI Serving Architecture
Designing systems that serve ML models at scale
About this session
A 60-minute live deep-dive into the architectures that power modern AI-serving systems: request batching, feature stores, model routing, cold-start mitigation, and latency SLOs. We'll walk through real-world reference designs and common interview framings so you can reason about tradeoffs on the spot.
Lecture: author of showoffer
Tony
Engineer Manager
Engineering Manager: Over 12 years of industry experience with 6 years in engineering leadership, building and scaling tech infrastructure teams that deliver end-to-end large-scale distributed systems.
Reviews
Reviewer 19N9
Apr 23, 2026The trade-offs around tensor vs pipeline vs expert parallelism were explained clearly with the failure modes of each. Recommend.
Reviewer 1F2I
Apr 22, 2026The breakdown of the ChatGPT-style serving stack — from request batching to KV-cache reuse — was the clearest I've seen. The coach made GPU economics easy to reason about. Highly recommend.
Reviewer 0EFA
Apr 22, 2026Loved the focus on tail-latency at the token level. Most resources stop at request-level metrics; this session went deeper and it shows.
Reviewer 0YH8
Apr 22, 2026Strong reference solution for the streaming response variant. Walking through SSE handling, backpressure, and client reconnects was very interview-relevant.
Reviewer 0ZOO
Apr 22, 2026Excellent treatment of the scheduling problem — how to mix prompt-heavy and generation-heavy requests on the same fleet. Practical and intuitive.
Reviewer 1VRZ
Apr 22, 2026Great session. The discussion on continuous batching vs static batching and when each makes sense was directly applicable to a real interview question I got last week.
Reviewer 0GRI
Apr 22, 2026The coach's framing of 'compute-bound vs memory-bound' phases of inference is now my go-to mental model. Recommend for anyone targeting AI infra roles.
Reviewer 1V6P
Apr 22, 2026Genuinely the first time I understood why speculative decoding matters and where it breaks. Worth attending live for the Q&A alone.
Reviewer 0IPF
Apr 22, 2026The deep-dive on KV-cache memory pressure and how it bounds concurrency was the most useful 15 minutes of any system design session I've taken.
Reviewer 1H5J
Apr 22, 2026Came away with a real framework for capacity-planning an LLM serving cluster. The numbers exercises were very useful.
Pre-session Q&A
Questions confirmed attendees asked before this session.
No questions yet.