NEW: ML Mock & Coaching now available

Questions

CI/CD System

OpenAI

Design a CI/CD System - This system design covers workflow orchestration, parallel execution, artifact management, and real-time visibility at scale.

50 min read

Staff+ Engineer: 10+ years of experience in build systems and CI/CD pipelines. Tech lead for distributed workflow orchestration at scale. Engineering Manager: Over 12 years of industry experience with 6 years in engineering leadership, building and scaling tech infrastructure teams that deliver end-to-end large-scale distributed systems.

Challenge

Think Beyond the Happy Path

System design interviews at companies like OpenAI and Meta often go beyond textbook answers—they test how you reason under constraints, design for scale, and handle failure gracefully.

Here are just a few of the questions you might face in CI/CD design:

  • Concurrency & Queuing: When multiple developers push code at once, how would you orchestrate and prioritize builds to optimize resource usage and minimize delays?
  • Multi-Tenant Architecture: If multiple teams share the same CI/CD platform, how would you isolate their pipelines, enforce fairness, and prevent noisy neighbor issues?
  • Immutable Build Artifacts: Would you enforce artifact immutability? If so, how would you ensure reproducibility and manage cache invalidation?
  • Rollback Safety: How would you design safe rollback mechanisms in a multi-service environment without risking cascading failures?

You don't need to answer all of these right away. A Senior/Staff+ Engineer doesn't stop at basic functionality — they anticipate edge cases, design for resilience, and push for production-grade reliability.

But great systems start with great questions. What would you ask next?

Problem Statement

Design a distributed CI/CD workflow system similar to GitHub Actions that automatically triggers and manages predefined workflows in response to repository events. The system should:

  1. Trigger workflows on each Git push using the repo-defined workflow file.
  2. Run workflows in isolated sandboxes with scoped secrets.
  3. Provide real-time visibility into execution (status, step progress, logs). And the scale is enterprise level with 100,000+ concurrent workflows.
What is CI/CD?

CI/CD isn't just another tech buzzword - it's the backbone of modern software development that keeps your team shipping fast without breaking things.

CI (Continuous Integration) - Your Code's Quality Guardian

Think of CI as that super diligent colleague who checks everyone's work before it gets merged. Every time someone pushes code:

  • Code gets automatically integrated into the main branch
  • Tests run immediately to catch bugs before they spread
  • Fast feedback loops mean developers know within minutes if they broke something

CD (Continuous Delivery/Deployment) - Your Release Automation Hero

  • Continuous Delivery keeps your code in a "ready-to-ship" state at all times. Think of it as you are dressed up and ready for the party, but you still decide when to actually go.
  • Continuous Deployment takes it a step further - every change that passes your tests automatically goes to production. It's like having a fully automated assembly line that just keeps churning out releases.

Functional Requirements

FR1 – Trigger workflows on each Git push using the repo-defined workflow file

This covers event ingestion, authentication, deduplication, workflow parsing, and DAG construction from the repository's workflow configuration.

FR2 – Run workflows in isolated sandboxes with scoped secrets

Jobs must execute in containerized or VM-based isolation with per-job resource allocation, secret injection, and artifact sharing between tasks.

FR3 – Provide real-time visibility into execution (status, step progress, logs)

Users should see live log streaming, step-by-step progress, and status updates as workflows execute across distributed runners.

Non-Functional Requirements

Non-Function Metrics Assumptions

Let's assume the system handles ~10 pushes per second on average, and each workflow takes about 15 minutes to run. That translates to roughly 10,000 concurrent workflows under normal load. During peak traffic — such as coordinated releases or mass test retries — the system may need to handle up to 10x the baseline, i.e., 100,000 concurrent workflows. This framing helps us size the CI/CD platform not just for today's team size but for realistic workload bursts, ensuring fast feedback, minimal queuing, and uninterrupted developer productivity.

NFR1 – High Concurrency

The system shall support concurrent execution of at least 100,000 jobs across the cluster.

NFR2 – Job Isolation

The system shall enforce per-job container isolation to limit blast radius and prevent cross-workflow interference.

NFR3 – Multi-Tenant Fairness

The system shall enforce per-organization/repo quotas to guarantee fairness and prevent noisy-neighbor effects.

NFR4 – Durability & Recovery

The system shall persist workflow state and artifacts across failures to enable safe recovery.

NFR5 – Low Latency Trigger Processing

The system shall process workflow triggers within 10 seconds of receiving a Git event during peak time.

Requirement Summary

Functional Requirements (FRs)
NameDescription
1. Workflow TriggeringTrigger workflows on Git push events using repo-defined YAML configs with DAG construction.
2. Isolated ExecutionRun jobs in containerized sandboxes with scoped secrets and artifact sharing.
3. Real-Time VisibilityStream logs live, show step progress, and push status updates to users.
Non-Functional Requirements (NFRs)
NameDescription
1. High ConcurrencySupport 100K+ concurrent jobs across the cluster.
2. Job IsolationContainer-level isolation to prevent cross-workflow interference.
3. Multi-Tenant FairnessPer-org/repo quotas to prevent noisy-neighbor effects.
4. Durability & RecoveryPersist state and artifacts for crash recovery.
5. Low Latency TriggersProcess triggers within 10s during peak load.

Core Entities

Core Entities
NameDescription
RepositoryA registered source code repository within the CI/CD system.
Workflow_RunA versioned specification tied to a commit that defines the pipeline.
TaskA single node within the workflow's DAG, executed in isolation.
Task_AttemptAn individual execution of a task, tracking retries and logs.
ArtifactFile (build output, logs) produced by a task or task_attempt.
Entity Relationships

GitHub Event → WorkflowRun (one per trigger) → Tasks (one per task, DAG edges from prerequisites[]) → TaskAttempts (one per execution try).

  • Repository — The origin of workflow definitions and the source of truth for pipeline execution.
  • Workflow_Run — Declares Tasks, their dependencies, and the steps each job must execute.
  • Task — Each task is executed in its own isolated environment (container, VM, or sandbox) to ensure reproducibility and fault isolation.
  • Task_Attempt — Represents a concrete execution, ordered and tracked for retries, logging, and observability.
  • Artifact — Build outputs that flow between tasks, stored immutably for reproducibility.

High Level Design

Tip

How to Deliver High-Level Design in Your Interviews?

Many candidates worry that if they don't mention everything upfront, they'll miss key evaluation points. But rather than frontloading everything, it's better to start from flow, show awareness of interesting decision points, and invite the interviewer to go deeper later.

Here's a structured 5-step approach:

  1. Truncate Each Functional Requirement into Building Blocks — Break each FR into clear, sequential building blocks (e.g., "On Git push → fetch YAML → parse DAG → enqueue jobs").
  2. Apply Core Design Principles per Block — Ask: Should it be sync or async? What are the durability guarantees? Do we need deduplication, retries, or fallbacks?
  3. Derive Entities from Flow, Not the Other Way Around — Identify which objects need to persist across blocks, then define entities.
  4. Draw a Text or Component Diagram to Tie It Together — Show ownership and lifecycle clearly.
  5. Circle Back to Deep Dives as Needed — Set expectations for later discussion.

FR1 - Trigger workflows on each Git push

In a modern CI/CD system, every pipeline is triggered by a Git event — such as a push, pull request, or merge. To manage this reliably at scale, we introduce an Ingress Service as the entry point.

Core Responsibilities:

  • Event ingestion & authentication: Securely receive and validate webhook events from Git providers.
  • Dedup: Git providers use at-least-once delivery, so events may be resent due to retries or network issues.
  • Workflow retrieval: Fetch the workflow definition pinned to the referenced commit SHA.
  • Validation: Ensure configuration correctness, syntax validity, and policy compliance.
  • Execution planning: Build a DAG of jobs and dependencies, then persist the run plan in durable storage.

Key Design Principles:

  • Asynchronous-first: Acknowledge webhooks quickly; defer heavy lifting off the hot path.
  • Exactly-once delivery: Deduplicate at ingress, preserve immutability, and ensure events are consumed reliably.
  • Durable state persistence: Persist workflow state and artifacts so execution can survive crashes.
cicd-FR1

Click to expand

Let's walk through how this works step by step.

Step 1: GitHub Push Event

GitHub Actions is powered internally by GitHub's event system, which is built on webhooks. When something happens (push, pull request, issue comment, release, etc.), GitHub fires an internal webhook-like event.

Step 2: Authenticate GitHub Action Event

The first responsibility is to authenticate the event via API Gateway, since processing an unauthenticated or replayed request could allow attackers to trigger arbitrary builds.

How are GitHub webhook events authenticated?

We expose a public endpoint (POST /webhooks/git) where the handler auto-detects the Git provider based on headers. Authentication varies by provider:

Authentication Methods by Provider
NameDescription
GitHubHMAC-SHA256 via X-Hub-Signature-256
GitLabShared token via X-Gitlab-Token
BitbucketJWT or shared HMAC (if configured)

All signatures are verified against the raw request body using constant-time comparison to mitigate timing attacks. We also support dual-secret rotation so keys can be rotated without downtime.

HMAC (Hash-based Message Authentication Code) is a cryptographic technique that uses a secret shared key and a hash function to verify both the integrity and authenticity of a message.

Step 3: Deduplication

Git providers use at-least-once delivery, meaning the same event may be sent multiple times due to retries, delays, or transient network failures. Without protection, this can cause duplicate workflow runs. We'll cover this in DD1.

Step 4: Fetch & Validate Workflow Definition

Once the event is authenticated and deduplicated, we fetch the workflow definition (e.g., .ci/workflows.yaml) pinned to the exact commit SHA referenced in the event.

Tip

Fetching by commit SHA, rather than by branch name (main), ensures that the workflow definition is tied to an immutable snapshot of the codebase. This makes builds reproducible and guards against "moving target" definitions that could change between retries.

Step 5: Parse and Validate the Workflow

After fetching, the workflow YAML is parsed into an internal representation. Validation is multi-layered:

  1. Schema validation: fields, formats, types.
  2. Policy linting: enforce org rules.
  3. Semantic validation: detect cycles in the prerequisites[], downstream[] graph, invalid job references, undeclared artifact names.
Tip

Stricter validation may block valid edge cases, but it reduces runtime failures and strengthens security. Early failure provides clear feedback to developers and saves cluster resources.

Step 6: Construct DAG

Once validated, the workflow definition is transformed into a Directed Acyclic Graph (DAG) of tasks.

DAG vs. Step-by-Step Workflow

In a CI/CD system, you can either model execution as a DAG or as a Step-by-Step flow.

A DAG gives you parallelism, the ability to fan out and fan in jobs, handle matrices, and retry failed nodes independently — which makes pipelines faster and more resilient, though it comes at the cost of added complexity in scheduling and state management.

A Step-by-Step flow is much simpler to reason about: each stage runs after the previous one, which works well for short, strictly linear pipelines, but it limits scalability since failures often require re-running the entire pipeline and there's no room for parallelism.

We suggest defaulting to a DAG for parallelism and selective retries, but collapsing to Step-by-Step when the pipeline is strictly linear.

The execution model is layered:

1. Workflow_run: Provides a reliable top-level object for retries, history, and observability.

  • Captures the entire triggered workflow as a durable record.
  • Links together the GitHub event, commit SHA, workflow file, and run status.

2. Tasks: Breaking runs into tasks exposes parallelism, enforces dependencies, and isolates execution.

  • Each task in the YAML becomes a Task node in the DAG.
  • Dependencies are modeled as edges; matrix strategies expand into multiple task variants.

3. Task_Attempts: Attempts make transitions idempotent and auditable — retries don't corrupt prior state.

  • Every actual execution is a Task_Attempt row, recording worker ID, start/end times, exit codes, and logs.
  • Retries create new attempts instead of overwriting history.
Tip

Durable Persistence

The DAG is persisted in a durable store, so if the controller crashes, a reconciler can rehydrate state and resume execution. This persistence adds minor latency and storage cost, but in CI/CD systems, durability and correctness far outweigh raw speed.

Sign in to continue reading

"CI/CD System" requires a free account to access.

Sign in to continue