Early access: your sandbox is free, with $5 of AQ Composer credits every month. Your own subscriptions stay unmetered. Start free

aq.dev / guides / coding-agent-sandboxes-explained

Coding Agent Sandboxes, Explained

A coding agent sandbox is an isolated execution environment where an AI coding agent like Claude Code or Codex can read files, install packages, and run commands without being able to touch the rest of your machine, your credentials, or your network. In late September 2026 the idea went mainstream in one week: DigitalOcean launched Managed Agents into public preview on September 22, running each agent session in its own Firecracker microVM, and Docker launched Cloud Sandboxes on September 24, moving its microVM sandboxes onto hosted compute. Both follow Cloudflare open sourcing Cloudflare OS in August 2026 with a sandboxed code runtime at its core. This guide defines the terms, compares what each option isolates, persists, and bills, and covers the questions a sandbox alone does not answer.

What is a coding agent sandbox?

An agent is only useful when it can execute: run the test suite, install a dependency, start the dev server, use git. Every one of those is an arbitrary shell command, and an agent that can run arbitrary commands can also delete files outside the project, read SSH keys and secrets, or leak them to a server it was tricked into contacting through prompt injection. The traditional answer is a permission prompt per command, which is safe and exhausting; permission modes tune that tradeoff. A sandbox changes the deal: instead of policing each command, you bound what any command can reach, which is what makes unattended and overnight runs defensible, and one half of limiting what a coding agent can destroy.

Four terms, defined in one sentence each

  • Container. An isolated process group sharing the host's kernel, separated by Linux namespaces and cgroups: lightweight and fast, but a kernel exploit escapes it.
  • MicroVM. A minimal virtual machine with its own guest kernel under hardware virtualization (Firecracker, the open-source AWS project behind Lambda, boots one in roughly 125 milliseconds), giving VM-grade isolation at close to container-grade startup.
  • Virtual machine (VM). A full emulated computer with its own kernel and operating system: the strongest practical boundary, and the slowest to create.
  • Git worktree. A second working directory of the same repository on its own branch: it isolates an agent's code changes from yours, but not its execution, so worktrees and sandboxes solve different halves of the problem.

The isolation spectrum

Sandboxing is not one technique. The options run from policy to hardware, and vendors sit at different points:

OS-level policy. The lightest boundary restricts what an ordinary process can touch. Claude Code ships a built-in sandbox for its Bash tool, as of October 2026 built on macOS Seatbelt and Linux bubblewrap: you declare which files and network domains commands may reach, the OS enforces it for every command and child process, and Anthropic publishes the mechanism as an open-source npm package (sandbox-runtime). No separate machine, no extra cost, and the host's kernel is the boundary.

Containers. One step up, the agent runs in a container with only the project mounted in. This blocks casual filesystem damage, but agents that themselves need Docker get awkward, and the shared kernel is why security-sensitive platforms moved past it (some interpose a userspace kernel, gVisor, to harden it).

MicroVMs. The pattern the September launches share. Docker Sandboxes (the local, free version, with microVM isolation since January 2026 on macOS and Windows) give each agent a microVM with its own kernel and Docker daemon, with only your project workspace mounted in. DigitalOcean's Harness Runtime creates a dedicated Firecracker microVM per session. Hardware virtualization means even a kernel exploit inside the sandbox stays inside the sandbox.

Full VMs. The boundary teams already trust for everything else: coarser-grained than a microVM per task, but a VM persists for months, holds real credentials safely, and behaves like the machines production software runs on. SDK-first services such as E2B and Modal sit between these last two points, renting per-second microVMs to developers building agent products.

The new sandboxes, compared (October 2026)

Docker Cloud Sandboxes (launched September 24, 2026) take the local microVM sandbox and run it on Docker-managed compute, so an agent keeps working after the laptop closes. The CLI is the same sbx tool, and one command (sbx move, with the --to cloud flag) captures a sandbox's filesystem and recreates it on the other side, in either direction. Pre-built Kits, now standard OCI images, package agents including Claude Code, Codex, Copilot, and OpenCode, and inject secrets through a proxy so the agent never sees the actual values. Pricing as of October 2026 is metered by the second, from $0.07 per hour (1 vCPU, 2 GiB) to $1.12 per hour (16 vCPU, 32 GiB), with volumes and egress free; sessions default to one hour and cap at 24 hours.

DigitalOcean Managed Agents (public preview since September 22, 2026) runs each session of Claude Code, Codex CLI, OpenCode, Hermes, or LangGraph (or your own agent as an OCI image) in a dedicated Firecracker microVM. Its distinctive move is economic: a session auto-pauses when idle, defined as no outgoing LLM or tool calls, resumes in about 305 milliseconds by DigitalOcean's measurement, and stops CPU and memory charges while paused. Its Action Gateway brokers credentials for external tools at execution time so they never reach the model or the sandbox. DigitalOcean's announcement lists active-usage pricing of $0.044 per vCPU-hour and $0.0095 per GB-hour.

Cloudflare OS (open sourced August 5, 2026, Apache 2.0) is a different animal: not a dev machine in the cloud but an agent workspace for every employee, where the agent writes and executes JavaScript or TypeScript in a sandboxed V8 isolate runtime inside a per-workspace Durable Object. There is no shell, no package manager, and no git in the sandbox; our Cloudflare OS explainer covers what that model is and is not good for.

OptionWhat it isolatesWhat persistsBilling (as of October 2026)
Claude Code sandbox (built in)Filesystem paths and network domains, via OS policy on your own machineEverything (it is your machine)Free with Claude Code
Docker Sandboxes (local)MicroVM with own kernel and Docker daemon; project workspace mounted inSandbox filesystem on your machineFree
Docker Cloud SandboxesSame microVM, on Docker-managed computeFilesystem state; moves laptop to cloud and back via sbx movePer second, $0.07 to $1.12 per hour by size
DigitalOcean Managed AgentsDedicated Firecracker microVM per session, own filesystemSession state across pause, resume, devices, and teammatesActive usage: $0.044 per vCPU-hour, $0.0095 per GB-hour; paused sessions free of CPU and memory charges
Cloudflare OS (Code Mode)V8 isolate inside a per-workspace Durable Object; no shell or gitWorkspace state and files in the Durable ObjectOpen source; runs on your Cloudflare account's Workers usage

What a sandbox does not answer

A sandbox bounds one agent's blast radius for one run. Teams adopting agents hit questions that sit outside that boundary:

  • Who can see the session? An agent working unattended in a cloud microVM is invisible by default. When it stalls or needs a decision, someone has to notice, open the session, and steer it.
  • How do two people share one session? Handoffs need the session itself to be a shared, addressable thing, not a process inside a box only one laptop can reach.
  • What did the agent actually do? Isolation during the run does not tell you afterwards whether the work is right. The record of commands run and output produced is what makes five overnight agents reviewable the next morning.
  • Where do real credentials live? The platforms are converging on brokering (Docker's Kits proxy secrets, DigitalOcean's Action Gateway injects them at execution time), but an agent that opens pull requests still needs to act as a specific person, with that person's git identity and review accountability.
  • How do parallel agents share a repository? Five sandboxes are five isolated machines, but pointing them at the same branch recreates the collision problem; per-task git worktrees on separate branches are the complement sandboxes do not include.
  • Can anyone see the app? A dev server inside a sandbox is only useful when a teammate, a designer, or a PM can open it in a browser and react.

Where AQ fits

AQ is the multiplayer coding harness where engineering teams run AI coding agents like Claude Code and Codex together: shared live terminals, a code editor, and app previews, in your own cloud. In this guide's terms, AQ takes the full-VM point on the spectrum and builds the team layer the sandbox launches leave open. Agents run as real CLIs (Claude Code, Codex, Cursor Agent, Kimi, Grok, or plain shells) in persistent tmux sessions on your team's VM, either machines you connect from your own cloud or a dedicated always-on AQ-managed VM in its own isolated network, with no shared multi-tenant execution tier. Each workspace gets its own isolated git worktree on its own branch, with dependencies installed automatically, so parallel agents never collide.

The team questions above are the product. Every session streams live to the browser, so teammates open the same workspace and watch the same agent work; typing into someone else's terminal is delegated by its owner approving a control request in one click, and sessions survive a closed laptop and resume from any device. Each engineer logs into the CLIs with their own Claude or OpenAI account (AQ never marks up usage on your own subscriptions), commits and PRs use per-user GitHub auth, and every workspace gets a live dev-server preview with shareable links that work without an account for viewing. The Free plan is itself a sandbox in this guide's sense: a personal machine AQ creates in an isolated network, nothing to install and no time limit. The Team plan is $50 per user per month in early access (standard $200, billed monthly), with the rate locked for your first 12 months.

Plainly: to bound one agent's blast radius on your laptop, Claude Code's built-in sandbox or local Docker Sandboxes already do that well and cost nothing, and the cloud sandbox platforms earn their keep when runs must outlast the laptop. AQ is for the step after that, when the point is not just that agents run somewhere safe, but that your team can watch, steer, share, and review them running together.

Frequently asked questions

Do I need a sandbox to run Claude Code or Codex safely?

For supervised work with permission prompts on, many developers run without one. A sandbox becomes important the moment you grant autonomy: skipping per-command approval, running unattended, or running overnight. As of October 2026 the zero-cost starting points are Claude Code's built-in OS-level sandbox (macOS Seatbelt, Linux bubblewrap) and Docker's free local Sandboxes, which give each agent a microVM with its own kernel.

What is the difference between a sandbox and a git worktree?

A sandbox isolates execution: what files, network, and system resources an agent's commands can reach. A git worktree isolates changes: a second working directory of the same repository on its own branch, so two tasks never edit the same checkout. They are complements, not substitutes. Parallel agents want both: a bounded place to run, and a separate worktree each so their diffs do not collide.

Is a Docker container enough to sandbox a coding agent?

It blocks casual damage, but containers share the host's kernel, so a kernel exploit escapes the boundary, and agents that need to run Docker themselves get awkward inside a container. That is why the platforms that launched in September 2026 (Docker Cloud Sandboxes, DigitalOcean Managed Agents) both chose microVMs, which carry their own guest kernel under hardware virtualization while starting in well under a second.

What does it cost to run a coding agent in a cloud sandbox?

As of October 2026, Docker Cloud Sandboxes meter by the second from $0.07 per hour for a 1 vCPU, 2 GiB sandbox up to $1.12 per hour for 16 vCPU, 32 GiB, with volumes and egress free. DigitalOcean Managed Agents bills active usage at $0.044 per vCPU-hour and $0.0095 per GB-hour, and stops CPU and memory charges while a session is auto-paused waiting on nothing. Idle time is the number to watch: an agent session spends much of its life waiting on the model, so pause behavior and per-second metering matter more than the headline rate.

Can my team see what an agent is doing inside a sandbox?

Mostly no, and that is the gap to plan for. A sandbox is an isolation boundary, not a collaboration surface: the session lives inside a box that one CLI or one account controls, and a teammate cannot watch it, take it over, or review afterwards what commands it ran. Platforms differ (DigitalOcean persists session state across devices and teammates; Docker's sbx CLI addresses sandboxes by name), but if the goal is several people driving and reviewing shared agent sessions, that is a workspace layer on top of the sandbox, which is the layer AQ provides.

Are cloud sandboxes the same as cloud development environments like Codespaces?

They rhyme but optimize differently. A cloud development environment is sized and billed for a human working an eight-hour day, with an editor attached. An agent sandbox is built for machine-driven, bursty, often parallel work: created per task in seconds, strongly isolated because the occupant is untrusted, and billed per second because fleets of them come and go. The September 2026 launches are the agent-shaped version of infrastructure that CDEs built for people.