Sandboxing an AI agent means running the code and commands it generates in an isolated environment with no standing credentials, default-deny network egress, hard resource limits and a complete audit log, so a manipulated or mistaken instruction cannot reach the rest of your infrastructure. Ordinary containers share the host kernel. For untrusted agent code, a user-space kernel such as gVisor or a microVM such as Firecracker or Kata Containers gives a stronger boundary.
Agent-written code is untrusted input
When an agent runs code, a model wrote that code while being steered by a prompt, retrieved documents and earlier tool outputs. Any of those can carry instructions the user never wrote. In an agent with an execution tool, that indirect prompt injection turns into commands on your hardware. The user can be trustworthy and the model well aligned while the code is still hostile.
Anthropic published a plain example in May 2026. During an internal red-team exercise, a phishing email hid instructions to read a local cloud credentials file and post it to an outside endpoint. Claude Code completed the exfiltration in 24 of 25 retries. Anthropic’s conclusion was that only the environment holds in that situation: egress controls, and filesystem boundaries that keep the credentials out of reach.
OWASP’s Top 10 for Agentic Applications, published in December 2025, names the risk ASI05, Unexpected Code Execution. Treat every execution as a stranger’s script that may try to read secrets, reach the network or leave something behind.
Isolation options compared
| Option | Boundary | What an escape needs | Trade-offs | Good fit |
|---|---|---|---|---|
| OS-level sandbox (seccomp, Landlock, bubblewrap, Seatbelt) | Kernel-enforced limits on one process | A kernel bug or a gap in the policy | Light and fast; policies are easy to get wrong; shares the host kernel | Coding-agent CLIs on developer machines |
| Standard container (runc) | Namespaces and cgroups on a shared kernel | A kernel or container-runtime vulnerability | Familiar and fast; weakest boundary for hostile code | Reviewed, trusted code |
| gVisor | A user-space application kernel that handles the container’s system calls | Breaking gVisor, then the host kernel | Not every syscall or /proc file is implemented; syscall- and file-heavy work runs slower | Analysis scripts and most agent code |
| MicroVM (Firecracker, Kata Containers) | Hardware virtualisation with a minimal virtual machine monitor | A hypervisor escape | Needs KVM on the host; more moving parts; strongest practical boundary | Builds, third-party packages, multi-tenant execution |
| Dedicated VM or host | A separate machine | Compromising the management plane | Slow and costly per session | Privileged work under approval |
gVisor is the layer Anthropic says isolates code execution in claude.ai. Firecracker, which powers AWS Lambda, states boot times under 125 ms and under 5 MiB of memory overhead per microVM. Kata Containers runs each pod in a lightweight VM behind the standard container interfaces.
The shared kernel is not a theoretical weakness. CVE-2024-21626, known as Leaky Vessels, let a malicious image reach the host filesystem through runc versions before 1.1.12. Three more runc flaws disclosed in November 2025 allowed escapes to the host through race conditions with shared mounts. NIST’s container security guide, SP 800-190, advised back in 2017 that only containers of the same purpose, sensitivity and threat posture should share a host kernel. Whatever tier you choose, patch the runtime on a schedule and turn on user namespaces, which runc’s maintainers recommend as the main mitigation for the 2025 flaws and which Kubernetes made generally available in v1.36.
The boundary is only half of a sandbox
Kernel isolation decides how hard an escape is. These controls decide how much a manipulated script can achieve without escaping at all:
- No standing credentials. Mount no SSH keys, cloud CLIs, kubeconfigs or service tokens. When a task needs access, a broker injects a short-lived, task-scoped credential. See the identity and containment checklist.
- Default-deny egress. Kubernetes NetworkPolicy can deny all egress, but only if your network plugin enforces it, and Firecracker filters no traffic itself. Keep allowlists short: in Anthropic’s write-up, a sandbox held while files still left through an allowlisted API domain, using an attacker’s own API key. Serve packages from an internal mirror, and test the controls as the egress-control guide describes.
- Ephemeral state. A fresh filesystem per session, destroyed afterwards, with no volumes shared across users or tenants.
- Hard limits. CPU, memory, disk, process count and wall-clock time, so a runaway loop dies instead of starving the node.
- Untrusted output. Treat stdout and generated files as data, not instructions, before another agent or a user consumes them.
- A complete record. Log every command, exit code and file written, tied to the agent, user and workflow that triggered it.
Choosing a tier by what the agent does
- Calculations on data the user can already see: gVisor or a microVM, with no network.
- Installing packages, building or testing code: a microVM, with egress only to internal mirrors.
- Shared platforms: one sandbox per session, never reused across users, with separate pools per tenant.
- Anything that changes production systems: not a sandbox problem. Use tools with scoped permissions and a human approval step for irreversible actions.
On Kubernetes, RuntimeClass lets one cluster run standard containers for platform services and gVisor or Kata for agent workloads. The SIG Apps agent-sandbox project, which reached v1.0 in August 2026, adds sandbox, template, claim and warm-pool resources on top and leaves the isolation itself to those runtimes.
Running sandboxes on-premises
Self-hosting adds duties of its own. MicroVMs need hardware virtualisation, so check whether your nodes are bare metal or VMs with nested virtualisation enabled. Keep warm pools of ready sandboxes to hide start-up time. Build sandbox images from a pinned internal registry, and mirror package repositories for air-gapped sites. Run sandboxes on their own node pool, so a build job never competes with inference for memory and a compromised sandbox sits away from model weights and retrieval indexes.
How VDF AI runs agent code
VDF AI’s sandboxed code execution tool runs agent-written code in an isolated, resource-limited sandbox scoped per tenant, and logs every command and file operation. The terminal tool runs only allow-listed commands, and tools are granted per role, so an agent receives execution only where its job needs it. The sandbox and workspace run in your data center or sovereign cloud, so source code, build artefacts and command output stay inside your perimeter. The wider code execution and workspace toolset covers tests, builds and Git under the same controls.
Sources and further reading
- Anthropic, How we contain Claude across products (May 2026)
- OWASP Top 10 for Agentic Applications for 2026
- gVisor documentation
- Firecracker
- Kata Containers
- Kubernetes documentation: RuntimeClass
- kubernetes-sigs/agent-sandbox
- runc advisory for CVE-2024-21626
- NIST SP 800-190, Application Container Security Guide
- Zero-trust network design for on-prem AI
Giving agents a code execution tool? Talk to us about running it inside your own infrastructure with isolation, egress control and a full audit trail.