NEW: ML Mock & Coaching now available

Thu, Jun 18, 20268:00 PM - 9:00 PM PDTSession ended

ML Model Distribution

P2P propagation, bandwidth math, and fault tolerance

About this session

A 60-minute live deep-dive into distributing a 500GB ML model from cloud storage to 1,000 GPU workers as fast as possible: why pipelines beat trees, chunking and manifests, saturating the network, integrity verification, and surviving worker failures mid-transfer. Great prep for AI-infra interviews at Anthropic and OpenAI.

Course Syllabus

Course Overview

Format:
Live Zoom, 90 minutes (75 min teaching + 15 min Q&A)
Audience:
Mid-level to Staff+ engineers prepping for AI-infra interviews (Anthropic, OpenAI)
Prerequisite:
Comfort with networking basics (bandwidth, throughput) and distributed systems

What You Will Learn

  1. Distribute a 500GB model to 1,000 GPU workers and minimize end-to-end deployment time — derive the math live.
  2. Explain why a pipeline beats a tree or sequential transfer, with numbers.
  3. Split the model into chunks and use a pre-distributed manifest (offsets, checksums) to verify integrity.
  4. Keep the system resilient when a worker fails mid-transfer, and detect when the job is done.
  5. Scale the same design from 1,000 to 10,000 workers.

What This Course Is NOT

  • Not model serving/inference — this is the distribution/delivery problem only
  • Not a memorization exercise — every number is derived live

Pre-Class Preparation (24 hours before)

Read the problem statement only (10 min):

Design a system to download a 500GB ML model from cloud storage and deliver it to all 1,000 GPU workers in a data center as fast as possible, with limited external bandwidth. The system must verify the integrity of every chunk against a pre-distributed manifest and keep working even if workers fail mid-transfer.

Think about (don't research yet):

  • In 30 seconds, what's your naive answer, and how long does it take?
  • Why split the file into small pieces, and how big should each chunk be?
  • What happens if a worker in the middle of a chain breaks?

Come to class with a number written down. Wrong answers are the most valuable starting point.

Lecture: author of showoffer

Tony

ShowOffer Coach

Engineering Manager: Over 12 years of industry experience with 6 years in engineering leadership, building and scaling tech infrastructure teams that deliver end-to-end large-scale distributed systems.

Reviews

5.0(1)
  • Reviewer 013R

    Jun 23, 2026

    I couldn't attend live but watch the video afterward. A quick question on the math. It seems that the design assume that at anyone tick, one worker can send or receive from *one* other worker so that the 10Gpbs up and down links are not shared, correct?

Pre-session Q&A

Questions confirmed attendees asked before this session.

No questions yet.