NEW: ML Mock & Coaching now available

Tue, May 12, 20267:00 PM - 8:00 PM PDTSession ended

Job Scheduling Systems

CI/CD system + Sora

About this session

Master job scheduling and workflow orchestration by designing a CI/CD pipeline system and an AI video generation platform like Sora.

Course Syllabus

Course Overview

Format:
60-minute live session (50 min teaching + 10 min Q&A)
Audience:
Senior and Staff-level engineers preparing for system design interviews at companies like OpenAI, Anthropic, Google, Meta, Databricks

What You Will Learn

  1. Explain the architecture behind modern Job Scheduling Systems used in CI/CD and AI infrastructure
  2. Design a scalable workflow orchestration platform similar to GitHub Actions, Airflow, Temporal, or internal OpenAI pipelines
  3. Understand the difference between schedulers, orchestrators, executors, and task queues
  4. Design DAG execution engines with retries, dependencies, fairness, and failure recovery
  5. Analyze how large-scale AI workloads (e.g. Sora video generation pipelines) are scheduled across GPU clusters
  6. Identify bottlenecks in distributed job systems: queue congestion, hot partitions, worker starvation, retry storms
  7. Justify design trade-offs between latency, throughput, durability, fairness, and cost
  8. Deliver a Staff-level system design answer with strong operational and scalability depth

What This Course Is NOT

  • Not “just draw queue + workers”
  • Not focused on memorizing architecture diagrams
  • Not purely academic — every design decision maps to production bottlenecks and operational incidents

Pre-Class Preparation (24 hours before)

Read the problem statement only (10 min):

Design a distributed job scheduling system that powers:

  • CI/CD pipelines for millions of builds per day
  • DAG-based workflow orchestration
  • GPU-intensive AI video generation pipelines similar to Sora

The system must support:

  • retries
  • priorities
  • dependency management
  • task queues
  • worker failures
  • autoscaling
  • fair scheduling across tenants
  • long-running workflows

Some jobs run for seconds, while others run for hours.

Think About (Don't Research Yet)

  • What breaks first at scale: database, queue, scheduler, or workers?
  • How would you prevent retry storms?
  • How would you resume a workflow after orchestrator failure?

Come to class with:

  • One naive architecture
  • One scaling concern
  • One failure mode you think is hardest to solve

Wrong answers are encouraged — that’s where the learning happens.

Lecture: author of showoffer

Jack

Staff+ Engineer

12+ years of experience in AI/ML infrastructure and distributed aggregation systems. Tech lead for fan-out architectures serving millions of product queries.

Reviews

5.0(1)
  • Reviewer 01UJ

    May 17, 2026

Pre-session Q&A

Questions confirmed attendees asked before this session.