ML Model Distribution
P2P propagation, bandwidth math, and fault tolerance
About this session
A 60-minute live deep-dive into distributing a 500GB ML model from cloud storage to 1,000 GPU workers as fast as possible: why pipelines beat trees, chunking and manifests, saturating the network, integrity verification, and surviving worker failures mid-transfer. Great prep for AI-infra interviews at Anthropic and OpenAI.
Course Syllabus
Course Overview
- Format:
- Live Zoom, 90 minutes (75 min teaching + 15 min Q&A)
- Audience:
- Mid-level to Staff+ engineers prepping for AI-infra interviews (Anthropic, OpenAI)
- Prerequisite:
- Comfort with networking basics (bandwidth, throughput) and distributed systems
What You Will Learn
- Distribute a 500GB model to 1,000 GPU workers and minimize end-to-end deployment time — derive the math live.
- Explain why a pipeline beats a tree or sequential transfer, with numbers.
- Split the model into chunks and use a pre-distributed manifest (offsets, checksums) to verify integrity.
- Keep the system resilient when a worker fails mid-transfer, and detect when the job is done.
- Scale the same design from 1,000 to 10,000 workers.
What This Course Is NOT
- Not model serving/inference — this is the distribution/delivery problem only
- Not a memorization exercise — every number is derived live
Pre-Class Preparation (24 hours before)
Read the problem statement only (10 min):
Design a system to download a 500GB ML model from cloud storage and deliver it to all 1,000 GPU workers in a data center as fast as possible, with limited external bandwidth. The system must verify the integrity of every chunk against a pre-distributed manifest and keep working even if workers fail mid-transfer.
Think about (don't research yet):
- In 30 seconds, what's your naive answer, and how long does it take?
- Why split the file into small pieces, and how big should each chunk be?
- What happens if a worker in the middle of a chain breaks?
Come to class with a number written down. Wrong answers are the most valuable starting point.
Lecture: author of showoffer
Tony
ShowOffer Coach
Engineering Manager: Over 12 years of industry experience with 6 years in engineering leadership, building and scaling tech infrastructure teams that deliver end-to-end large-scale distributed systems.
Reviews
Reviewer 013R
Jun 23, 2026I couldn't attend live but watch the video afterward. A quick question on the math. It seems that the design assume that at anyone tick, one worker can send or receive from *one* other worker so that the 10Gpbs up and down links are not shared, correct?
Pre-session Q&A
Questions confirmed attendees asked before this session.
No questions yet.