AI Infrastructure

Sandboxing AI Agents That Run Code: Containers, gVisor and MicroVMs in Private Infrastructure

An agent that writes and runs code is executing untrusted input on your infrastructure. How the isolation options compare, what a sandbox must control besides the kernel boundary, and how to run agent code execution safely inside an on-premises AI platform.

Sandboxing an AI agent means running the code and commands it generates in an isolated environment with no standing credentials, default-deny network egress, hard resource limits and a complete audit log, so a manipulated or mistaken instruction cannot reach the rest of your infrastructure. Ordinary containers share the host kernel. For untrusted agent code, a user-space kernel such as gVisor or a microVM such as Firecracker or Kata Containers gives a stronger boundary.

Agent-written code is untrusted input

When an agent runs code, a model wrote that code while being steered by a prompt, retrieved documents and earlier tool outputs. Any of those can carry instructions the user never wrote. In an agent with an execution tool, that indirect prompt injection turns into commands on your hardware. The user can be trustworthy and the model well aligned while the code is still hostile.

Anthropic published a plain example in May 2026. During an internal red-team exercise, a phishing email hid instructions to read a local cloud credentials file and post it to an outside endpoint. Claude Code completed the exfiltration in 24 of 25 retries. Anthropic’s conclusion was that only the environment holds in that situation: egress controls, and filesystem boundaries that keep the credentials out of reach.

OWASP’s Top 10 for Agentic Applications, published in December 2025, names the risk ASI05, Unexpected Code Execution. Treat every execution as a stranger’s script that may try to read secrets, reach the network or leave something behind.

Isolation options compared

OptionBoundaryWhat an escape needsTrade-offsGood fit
OS-level sandbox (seccomp, Landlock, bubblewrap, Seatbelt)Kernel-enforced limits on one processA kernel bug or a gap in the policyLight and fast; policies are easy to get wrong; shares the host kernelCoding-agent CLIs on developer machines
Standard container (runc)Namespaces and cgroups on a shared kernelA kernel or container-runtime vulnerabilityFamiliar and fast; weakest boundary for hostile codeReviewed, trusted code
gVisorA user-space application kernel that handles the container’s system callsBreaking gVisor, then the host kernelNot every syscall or /proc file is implemented; syscall- and file-heavy work runs slowerAnalysis scripts and most agent code
MicroVM (Firecracker, Kata Containers)Hardware virtualisation with a minimal virtual machine monitorA hypervisor escapeNeeds KVM on the host; more moving parts; strongest practical boundaryBuilds, third-party packages, multi-tenant execution
Dedicated VM or hostA separate machineCompromising the management planeSlow and costly per sessionPrivileged work under approval

gVisor is the layer Anthropic says isolates code execution in claude.ai. Firecracker, which powers AWS Lambda, states boot times under 125 ms and under 5 MiB of memory overhead per microVM. Kata Containers runs each pod in a lightweight VM behind the standard container interfaces.

The shared kernel is not a theoretical weakness. CVE-2024-21626, known as Leaky Vessels, let a malicious image reach the host filesystem through runc versions before 1.1.12. Three more runc flaws disclosed in November 2025 allowed escapes to the host through race conditions with shared mounts. NIST’s container security guide, SP 800-190, advised back in 2017 that only containers of the same purpose, sensitivity and threat posture should share a host kernel. Whatever tier you choose, patch the runtime on a schedule and turn on user namespaces, which runc’s maintainers recommend as the main mitigation for the 2025 flaws and which Kubernetes made generally available in v1.36.

The boundary is only half of a sandbox

Kernel isolation decides how hard an escape is. These controls decide how much a manipulated script can achieve without escaping at all:

  1. No standing credentials. Mount no SSH keys, cloud CLIs, kubeconfigs or service tokens. When a task needs access, a broker injects a short-lived, task-scoped credential. See the identity and containment checklist.
  2. Default-deny egress. Kubernetes NetworkPolicy can deny all egress, but only if your network plugin enforces it, and Firecracker filters no traffic itself. Keep allowlists short: in Anthropic’s write-up, a sandbox held while files still left through an allowlisted API domain, using an attacker’s own API key. Serve packages from an internal mirror, and test the controls as the egress-control guide describes.
  3. Ephemeral state. A fresh filesystem per session, destroyed afterwards, with no volumes shared across users or tenants.
  4. Hard limits. CPU, memory, disk, process count and wall-clock time, so a runaway loop dies instead of starving the node.
  5. Untrusted output. Treat stdout and generated files as data, not instructions, before another agent or a user consumes them.
  6. A complete record. Log every command, exit code and file written, tied to the agent, user and workflow that triggered it.

Choosing a tier by what the agent does

  • Calculations on data the user can already see: gVisor or a microVM, with no network.
  • Installing packages, building or testing code: a microVM, with egress only to internal mirrors.
  • Shared platforms: one sandbox per session, never reused across users, with separate pools per tenant.
  • Anything that changes production systems: not a sandbox problem. Use tools with scoped permissions and a human approval step for irreversible actions.

On Kubernetes, RuntimeClass lets one cluster run standard containers for platform services and gVisor or Kata for agent workloads. The SIG Apps agent-sandbox project, which reached v1.0 in August 2026, adds sandbox, template, claim and warm-pool resources on top and leaves the isolation itself to those runtimes.

Running sandboxes on-premises

Self-hosting adds duties of its own. MicroVMs need hardware virtualisation, so check whether your nodes are bare metal or VMs with nested virtualisation enabled. Keep warm pools of ready sandboxes to hide start-up time. Build sandbox images from a pinned internal registry, and mirror package repositories for air-gapped sites. Run sandboxes on their own node pool, so a build job never competes with inference for memory and a compromised sandbox sits away from model weights and retrieval indexes.

How VDF AI runs agent code

VDF AI’s sandboxed code execution tool runs agent-written code in an isolated, resource-limited sandbox scoped per tenant, and logs every command and file operation. The terminal tool runs only allow-listed commands, and tools are granted per role, so an agent receives execution only where its job needs it. The sandbox and workspace run in your data center or sovereign cloud, so source code, build artefacts and command output stay inside your perimeter. The wider code execution and workspace toolset covers tests, builds and Git under the same controls.

Sources and further reading


Giving agents a code execution tool? Talk to us about running it inside your own infrastructure with isolation, egress control and a full audit trail.

Frequently asked questions

Is a Docker container a secure sandbox for AI agents?

It is a useful layer, not a strong boundary for untrusted code. Containers share the host kernel, and runc flaws disclosed in 2024 and 2025 allowed escapes to the host. For code an agent writes, add a user-space kernel such as gVisor or use a microVM, and keep credentials and network access out of the sandbox.

Should I use gVisor or Firecracker for agent code execution?

gVisor runs as an OCI runtime and handles system calls in a user-space kernel, so it fits existing container tooling with little change; the cost is compatibility gaps and overhead on syscall-heavy work. Firecracker and Kata Containers use hardware virtualisation, which gives a stronger boundary but needs KVM on the host. Many teams use gVisor for routine execution and microVMs for builds and third-party code.

Should an agent sandbox have internet access?

Not by default. Default-deny egress blocks both data exfiltration and the download of arbitrary code. Where agents need packages, point them at an internal mirror of approved packages. If a task needs an external service, give the agent a governed tool with its own credentials and logging instead of opening the sandbox's network.

Does a sandbox replace human approval for agent actions?

No. A sandbox limits what code can reach while it runs. Approval governs actions whose effects leave the sandbox, such as merging code, changing records or sending messages. Use both: sandboxes for execution, and scoped tools with approval steps for anything irreversible.

Filed under
AI securityenterprise AI agentsAI infrastructureon-premises AIAI agent toolszero trust
On-Prem AI

Plan your on-prem AI deployment

Book an architecture call and we will scope a private, on-prem AI deployment for your environment — integrations, hardware, and governance included.

Keep reading