NEW: ML Mock & Coaching now available

Tue, Apr 21, 20267:00 PM - 8:00 PM PDTSession ended

AI Serving Architecture

Designing systems that serve ML models at scale

About this session

A 60-minute live deep-dive into the architectures that power modern AI-serving systems: request batching, feature stores, model routing, cold-start mitigation, and latency SLOs. We'll walk through real-world reference designs and common interview framings so you can reason about tradeoffs on the spot.

Lecture: author of showoffer

Tony

Engineer Manager

Engineering Manager: Over 12 years of industry experience with 6 years in engineering leadership, building and scaling tech infrastructure teams that deliver end-to-end large-scale distributed systems.

Reviews

5.0(10)
  • Reviewer 19N9

    Apr 23, 2026

    The trade-offs around tensor vs pipeline vs expert parallelism were explained clearly with the failure modes of each. Recommend.

  • Reviewer 1F2I

    Apr 22, 2026

    The breakdown of the ChatGPT-style serving stack — from request batching to KV-cache reuse — was the clearest I've seen. The coach made GPU economics easy to reason about. Highly recommend.

  • Reviewer 0EFA

    Apr 22, 2026

    Loved the focus on tail-latency at the token level. Most resources stop at request-level metrics; this session went deeper and it shows.

  • Reviewer 0YH8

    Apr 22, 2026

    Strong reference solution for the streaming response variant. Walking through SSE handling, backpressure, and client reconnects was very interview-relevant.

  • Reviewer 0ZOO

    Apr 22, 2026

    Excellent treatment of the scheduling problem — how to mix prompt-heavy and generation-heavy requests on the same fleet. Practical and intuitive.

  • Reviewer 1VRZ

    Apr 22, 2026

    Great session. The discussion on continuous batching vs static batching and when each makes sense was directly applicable to a real interview question I got last week.

  • Reviewer 0GRI

    Apr 22, 2026

    The coach's framing of 'compute-bound vs memory-bound' phases of inference is now my go-to mental model. Recommend for anyone targeting AI infra roles.

  • Reviewer 1V6P

    Apr 22, 2026

    Genuinely the first time I understood why speculative decoding matters and where it breaks. Worth attending live for the Q&A alone.

  • Reviewer 0IPF

    Apr 22, 2026

    The deep-dive on KV-cache memory pressure and how it bounds concurrency was the most useful 15 minutes of any system design session I've taken.

  • Reviewer 1H5J

    Apr 22, 2026

    Came away with a real framework for capacity-planning an LLM serving cluster. The numbers exercises were very useful.

Pre-session Q&A

Questions confirmed attendees asked before this session.

No questions yet.