Model Distribution
Model Distribution
About this session
Explore AI serving architecture through designing a ChatGPT-style inference pipeline and a model distribution system that fans large model weights out to a GPU fleet. Derive the BitTorrent-style propagation math live and stress-test the design at 10x scale.
Course Syllabus
Course Overview
- Format:
- 60-minute live session (50 min teaching + 10 min Q&A)
- Audience:
- Candidates preparing for system design interviews at companies like Anthropic, OpenAI, Meta, Google
- Prerequisite:
- Comfort with networking basics (bandwidth, TCP, HTTP), familiarity with distributed systems concepts (replication, consistency)
What You Will Learn
- From S3 to GPU HBM — what a 500 GB model really is.
- Distribute 500 GB to 1,000 GPUs in 7 minutes — derive the math live.
- Tracking, manifests, and the state that makes the pipeline debuggable.
- What breaks at 10,000 workers, and how to keep the system alive.
- A rack just died. 100 workers gone. What happens next?
- The killer perturbation: 500 GB uncompressed, 200 GB on the wire.
What This Course Is NOT
- Not a lecture — you will derive every number live
- Not a memorization exercise — you will be able to handle variants (e.g., 5 TB model, 10,000 workers, 1 Gbps pipe)
- Not architecture astronomy — every design choice ties back to a concrete bottleneck
Pre-Class Preparation (24 hours before)
Read the problem statement only (10 min):
Design a system to distribute a 500 GB ML model from S3 to 1,000 GPU workers in a data center. External bandwidth is 10 Gbps. Internal worker bandwidth is 10 Gbps full-duplex. Goal: minimize total time. Must survive worker failures.
Think about (don't research yet):
- In 30 seconds, what's your naive answer? How long does it take?
- What's the obvious bottleneck?
- Have you seen a similar pattern before?
Come to class with a number written down. Wrong answers are welcome — they're the most valuable starting point.
Lecture: author of showoffer
Tony
Engineer Manager
Engineering Manager: Over 12 years of industry experience with 6 years in engineering leadership, building and scaling tech infrastructure teams that deliver end-to-end large-scale distributed systems.
Reviews
Reviewer 09A5
Jun 1, 2026
Pre-session Q&A
Questions confirmed attendees asked before this session.
No questions yet.