Job Scheduling Systems
CI/CD system + Sora
About this session
Master job scheduling and workflow orchestration by designing a CI/CD pipeline system and an AI video generation platform like Sora.
Course Syllabus
Course Overview
- Format:
- 60-minute live session (50 min teaching + 10 min Q&A)
- Audience:
- Senior and Staff-level engineers preparing for system design interviews at companies like OpenAI, Anthropic, Google, Meta, Databricks
What You Will Learn
- Explain the architecture behind modern Job Scheduling Systems used in CI/CD and AI infrastructure
- Design a scalable workflow orchestration platform similar to GitHub Actions, Airflow, Temporal, or internal OpenAI pipelines
- Understand the difference between schedulers, orchestrators, executors, and task queues
- Design DAG execution engines with retries, dependencies, fairness, and failure recovery
- Analyze how large-scale AI workloads (e.g. Sora video generation pipelines) are scheduled across GPU clusters
- Identify bottlenecks in distributed job systems: queue congestion, hot partitions, worker starvation, retry storms
- Justify design trade-offs between latency, throughput, durability, fairness, and cost
- Deliver a Staff-level system design answer with strong operational and scalability depth
What This Course Is NOT
- Not “just draw queue + workers”
- Not focused on memorizing architecture diagrams
- Not purely academic — every design decision maps to production bottlenecks and operational incidents
Pre-Class Preparation (24 hours before)
Read the problem statement only (10 min):
Design a distributed job scheduling system that powers:
- CI/CD pipelines for millions of builds per day
- DAG-based workflow orchestration
- GPU-intensive AI video generation pipelines similar to Sora
The system must support:
- retries
- priorities
- dependency management
- task queues
- worker failures
- autoscaling
- fair scheduling across tenants
- long-running workflows
Some jobs run for seconds, while others run for hours.
Think About (Don't Research Yet)
- What breaks first at scale: database, queue, scheduler, or workers?
- How would you prevent retry storms?
- How would you resume a workflow after orchestrator failure?
Come to class with:
- One naive architecture
- One scaling concern
- One failure mode you think is hardest to solve
Wrong answers are encouraged — that’s where the learning happens.
Lecture: author of showoffer
Jack
Staff+ Engineer
12+ years of experience in AI/ML infrastructure and distributed aggregation systems. Tech lead for fan-out architectures serving millions of product queries.
Reviews
Reviewer 01UJ
May 17, 2026
Pre-session Q&A
Questions confirmed attendees asked before this session.