Think Beyond the Happy Path
Real production cloud IDEs aren't just about running code in a browser — they're about secure connectivity, burst-resistant provisioning, and reliable streaming under real-world network conditions.
Before diving into the material, take a moment to ask yourself:
- Do you know how to handle thousands of users opening the IDE at once during peak hours (class starts, workday mornings) without overwhelming VM provisioning or leaving users stuck in endless "starting" states?
- Do you know how to securely connect a browser to a remote devbox without exposing the VM to the public internet — while still achieving sub-second interactive latency?
- Do you know what happens when a browser disconnects for 30 seconds (Wi-Fi blip, laptop sleep, tab refresh) — how do you resume the stream without missing critical output or replaying too much?
- Do you know how to protect tail latency when one user produces massive stdout or spams execute requests — ensuring noisy neighbors don't degrade the experience for everyone else?
You don't need to answer all of these right away. A Senior/Staff+ Engineer doesn't stop at basic functionality — they anticipate edge cases, design for resilience, and push for production-grade reliability.
But great systems start with great questions. What would you ask next?
Problem Statement
What to Design?
Design a browser-based, cloud IDE (similar to Google Colab) where users can write code, run it on remote devbox, and view results in real time—without installing anything locally.
This question evaluates how you handle high scalability, resource scheduling and lifecycle management and low-latency streaming of logs/output at scale. The hardest parts are safely provisioning/tearing down execution environments and delivering interactive output efficiently and reliably.
What should be the correct design scope?
For scope alignment, we assume the system's responsibility ends at the control plane + transport layer:
- The system should accept an execute request, route it to the correct devbox, and stream outputs back to the browser.
- Once code is inside the devbox, the execution details are managed by the VM itself (process scheduling, language runtime behavior, terminal/kernel internals, etc.). In other words, we design the platform to provision/attach to environments, deliver execution requests, and stream output, but we do not implement an OS-level scheduler or a full kernel inside the VM—those are treated as implementation details of the devbox image.
To keep the interview scope bounded and the design concrete, we'll make 4 product-level assumptions:
-
Devbox is VM-based (by default).
We optimize for strong isolation and predictable resource accounting. Containers can be an implementation variant later, but the control-plane shape stays similar.
-
User ↔ workspace ↔ devbox is 1:1:1. A user has at most one active workspace in this system, and each workspace has at most one active devbox. This makes lifecycle, placement, and routing fully deterministic: user → workspace → devbox always resolves to a single target (no ambiguity, no multi-attach).
-
No devbox sharing.
The devbox is single-attached. We avoid multi-writer compute complexity (concurrent terminals/kernels, conflicting processes, shared state coordination). If we later support collaboration, it applies to WorkspaceFiles, not the live devbox.
-
Devbox disk persists via volume/snapshot.
Compute is replaceable, but disk is durable. If a VM restarts or is replaced, we reattach the same persistent volume (or restore from snapshot), so users don't lose their working directory or re-download dependencies.
Functional Requirements
FR1 — Workspace files (edit + persist)
- User can create/open a workspace and edit WorkspaceFiles in-browser
- Autosave with versioning
FR2 — Devbox provisioning & lifecycle (attach)
- Start/resume/suspend/stop a managed, time-bounded devbox (VM) for a workspace
- Devbox disk persists via volume/snapshot
FR3 — Execute commands (interactive)
- User can execute code/commands on the attached devbox
- User can cancel/interrupt the currently running command
FR4 — Real-time streaming (stdout/stderr/events)
- Stream stdout/stderr and execution events to the browser in real time
System Assumptions
- Scale: ~1M DAU, peak 200K concurrent IDE sessions (work/class hours).
- Execute requests (interactive): active session triggers ~1 execute / 3 min ⇒ 200K / 180 ≈ 1.1K exec/sec at peak; assume 10× burst (class starts/demos) ⇒ ~10K exec/sec peak.
- Streaming: ~50% sessions streaming at once ⇒ ~100K concurrent streams; assume ~1 KB/sec avg output ⇒ ~100 MB/sec aggregate throughput (plan 2–3× headroom).
- Devboxes (VM): default devbox ~1 vCPU + 2–4 GB RAM; workspace ↔ devbox is 1:1; keep 5–10% warm pool to reduce cold-start; auto-suspend after 10–15 min idle.
- Persistent disk: devbox uses a workspace-bound volume (or restore from latest snapshot) so files/deps persist across stop/start; compute is replaceable, disk is durable.
Non-Functional Requirements
NFR1 — Cluster Scalability & Fair Scheduling
Metric: peak devbox lifecycle requests (start/resume) 10K req/s
At peak hours, many users can open the IDE at the same time. The hardest scaling pressure is devbox start/resume (VM + disk attach), not the command execution inside the devbox. This NFR requires the system to handle large bursts of start/resume requests reliably and degrade predictably when capacity is tight.
Common deep dive interview question (team-scale runtimes)
- "During a global peak hour, a huge wave of users opens the IDE within minutes. How do you keep devbox start/resume working well at that scale, and what happens when there aren't enough resources?"
NFR2 — Secure Connectivity (Browser ↔ Remote Dev Env)
Metric: interactive connection setup p95 ≤ 1s (and stays encrypted end-to-end)
Users are coding in a browser, but the real execution happens in a powerful remote devbox. The key requirement is to provide a safe, reliable, encrypted channel from the browser to that devbox for terminal/execute/output streaming—without exposing the devbox directly to the public internet. This NFR forces the design to treat connectivity as a first-class problem: how sessions are established, how traffic is routed, and how we keep the devbox reachable only through controlled paths.
Common deep dive interview question (connectivity path)
- "How does the browser actually talk to the devbox? Describe the end-to-end connectivity flow for (1) attach/open IDE, (2) send an execute command, and (3) receive stdout/stderr/events. What protocol do you use and why?"
NFR3 — Real-time Output/Log Streaming at Scale
Metric: stream latency (devbox → browser) p95 ≤ 250ms
The IDE feels "local" when output shows up immediately: prints, errors, progress, terminal echo. The hard part isn't sending a few lines — it's staying smooth at high concurrency with real-world networks: tabs refresh, connections drop, and some sessions produce very noisy logs. This NFR pushes the design to handle reconnect, backfill (so users don't miss output), and flow control so streaming stays fast and stable under load.
Common deep dive interview question (reconnectivity)
- "A browser disconnects for ~30 seconds and then reconnects. How do you resume the stream so the user doesn't miss output, and you don't duplicate too much?"
Requirement Summary
| Functional Requirements (FRs) | |
|---|---|
| Name | Description |
1. Workspace files | User can create/open a workspace and edit WorkspaceFiles in-browser with autosave. |
2. Devbox provisioning & lifecycle | Start/resume/suspend/stop a managed, time-bounded devbox (VM) for a workspace. |
3. Execute commands | User can execute code/commands on the attached devbox and cancel/interrupt. |
4. Real-time streaming | Stream stdout/stderr and execution events to the browser in real time. |
| Non-Functional Requirements (NFRs) | |
|---|---|
| Name | Description |
1. Cluster Scalability | Peak devbox lifecycle requests (start/resume) 10K req/s. |
2. Secure Connectivity | Interactive connection setup p95 ≤ 1s, encrypted end-to-end. |
3. Real-time Streaming | Stream latency (devbox → browser) p95 ≤ 250ms. |
Core Entities
Core Entities from User Journey
A user works inside a Workspace (workspace_id, user_id, updated_at), which is the top-level container for files and settings. The editable content lives as WorkspaceFile (workspace_id, path, content_ref, version, updated_at) to support autosave and versioning.
When the user clicks "Connect," the system provisions a Devbox (devbox_id, workspace_id, state, vm_ref, last_active_at, expires_at) and attaches a durable DevboxDisk (volume_id, workspace_id, snapshot_ref, attached_devbox_id, state) so compute can be replaced without losing the working directory. Finally, to support real-time output plus refresh/reconnect, the client maintains a StreamCursor (workspace_id, devbox_id, stream_type, cursor, updated_at) so the streaming channel can resume from the last-seen offset and backfill gaps.
To keep things lightweight, we'll stick to the core entities implied by the journey (workspace → devbox → stream):
-
WorkspaceTop-level container the user works in (and the anchor for 1:1:1 mapping).
Example columns:
workspace_id,user_id,status,created_at,updated_at -
WorkspaceFile(aka "File")The editable content inside a workspace, with autosave + versioning.
Example columns:
file_id,workspace_id,path,latest_version,updated_at,content_ref -
Devbox(VM compute)The actual remote environment attached to the workspace.
Example columns:
devbox_id,workspace_id,state(STARTING/RUNNING/SUSPENDED/STOPPED),vm_ref,last_active_at,expires_at -
DevboxDisk(Volume/Snapshot)The durable disk for the devbox (compute replaceable, disk durable).
Example columns:
volume_id,workspace_id,attached_devbox_id,size_gb,snapshot_ref,state -
StreamCursorClient's last-seen offset for stdout/stderr/events so reconnect can backfill.
Example columns:
workspace_id,devbox_id,stream_type(STDOUT/STDERR/EVENT),cursor,updated_at
APIs
We can map the functional requirements back to a small, generic API surface, along with the entities. For brevity, we use one representative API per area and leave the obvious read/list variants for the audience to expand.
- FR1 — Workspace files:
POST /workspaces/{workspace_id}/files:write - FR2 — Devbox lifecycle:
POST /workspaces/{workspace_id}/devbox:transition(body hasdesired_state: START/RESUME/SUSPEND/STOP) - FR3 — Execute commands:
POST /workspaces/{workspace_id}/exec- returns
202 Accepted(or200) and streams output over the active stream connection
- returns
- FR4 — Streaming:
WS /streams
High-Level Design
To keep the High-Level Design easy to follow, we'll stay intentionally lightweight: for each functional requirement, we'll show (1) a step-by-step workflow, and (2) a diagram where one arrow maps to one step. This level of detail is usually enough to prove the system works end-to-end.
That said, we know the real interview signals often come from follow-up questions. So after HLD, we'll go deeper using the challenge-driven deep dives already implied by our NFRs (notebook collaboration consistency, runtime scaling and scheduling for large teams, environment lifecycle efficiency, secure client-to-runtime connectivity, and reliable log streaming with reconnect/backfill).