AI Engineering
tutorial
Featured

Your Agent Writes Code and Runs It Where Your Secrets Live

The feature that makes a code-writing agent useful, running code it generated a second ago, is also the widest hole in your stack. That code is written from input you do not control, and most teams run it in the same process or a plain container that can read every secret and call any host. Here is how to sandbox agent code execution properly: hard isolation, no credentials inside, a default-deny egress broker, strict resource limits, and a box you throw away after every task.

Viral Ruparel
11 min read
Share:

A team I worked with last quarter shipped a data-analysis agent that could write and run Python. A user uploads a spreadsheet, asks a question in plain English, and the agent writes pandas code, runs it, and hands back a chart. It demoed beautifully. It also ran that generated code in the same worker process as the API, with the production environment loaded, because that was the fastest way to get the demo working and nobody circled back.

The problem is not hypothetical and it is not exotic. One of their users uploaded a spreadsheet whose first cell contained a line of text asking the assistant to, as part of its analysis, read an environment variable and include it in the summary. The agent did exactly that. It wrote code that read os.environ, found the database URL and a payment key sitting right there, and dutifully printed them into the chat. No kernel exploit, no clever escape. The code did what the code said, and the code had been written from input the team did not control.

This is the shape of the whole problem. The single capability that makes a code-writing agent valuable, running code it generated a moment ago, is also the largest attack surface you will ever add to a stack. And the reflex to run that code in-process or in a plain container is how the surface turns into an incident.

The code your agent runs is written by someone you cannot see

Your own application code is a fixed artifact. You write it, you review it, you test it, and then it runs the same way a million times. You reason about it once.

Agent-generated code breaks every part of that. It is new on every turn. Nobody reviewed it. And it is shaped by whatever the agent read on the way to writing it: the uploaded file, the fetched page, the result of a tool call to some upstream system. If any of those inputs can be influenced by an attacker, and in a real product most of them can, then the person effectively authoring your runtime code is not the model and not you. It is whoever controlled that input.

So the correct default is the zero-trust one: treat every line of agent-generated code as if a hostile stranger wrote it, because sometimes one did. That sounds dramatic until you write down what the code can touch today, and realize the honest answer is usually "everything the process can."

Here is the naive version almost everyone starts with, so we are clear about what we are replacing.

# DO NOT DO THIS. Agent-authored code runs in your process,
# with your environment, your filesystem, your network.
def run_agent_code(code: str) -> str:
    local_ns: dict = {}
    # exec gives the generated code the full power of this interpreter:
    # os.environ, open(), socket, subprocess, the lot.
    exec(code, {"__builtins__": __builtins__}, local_ns)
    return str(local_ns.get("result", ""))

Every comment in that snippet is a door. os.environ is your secrets. open() is your filesystem. socket is your whole network. The agent does not need to be malicious to walk through one of those doors; it just needs to be convinced by its input that walking through is part of the task.

What a real sandbox actually enforces

A sandbox worth the name is not one control, it is five, and skipping any one of them tends to be the gap the incident walks through.

The first is hard isolation. A standard container shares the host kernel, which means the generated code's syscalls hit the same kernel running your other workloads, and a single kernel CVE is a host compromise. For untrusted code you want kernel-level isolation: a microVM (Firecracker and the managed services built on it) or a userspace kernel like gVisor. This is also why the ecosystem moved the way it did. OpenAI's Agents SDK added native sandboxed execution in 2026 with bring-your-own-sandbox providers like E2B, Modal, and Daytona, and Cloudflare shipped ephemeral Worker sandboxes. The common thread is that none of them run your agent's code on a shared kernel next to your app.

The second is no credentials inside the box. Not short-lived ones, not "scoped" ones, none. If a secret is reachable from inside the sandbox, a hostile program will reach it, and the whole premise is that the program might be hostile.

The third is default-deny egress. The sandbox gets no open outbound network. This single control neutralizes most exfiltration, because stolen data is only dangerous if it can leave.

The fourth is bounded resources: hard caps on CPU, memory, and wall-clock time, so a runaway loop or a crypto-miner planted by injected code dies in seconds instead of running up a cloud bill.

The fifth is a short life. The box is created for one task and destroyed after it. Persistent state across tasks is convenient and it is exactly how one user's poisoned run leaves something behind for the next user's run to pick up.

Here is the policy as a single object, so the guarantees are explicit rather than scattered across config files.

from dataclasses import dataclass, field

@dataclass(frozen=True)
class SandboxPolicy:
    cpu_cores: float = 1.0
    memory_mb: int = 512
    wall_clock_seconds: int = 20      # hard timeout; the box is killed after
    # Default-deny network. Only these hosts are reachable, and only
    # through the broker (see below), never directly from the sandbox.
    egress_allowlist: tuple[str, ...] = ()
    writable_paths: tuple[str, ...] = ("/tmp/work",)
    mount_secrets: bool = False       # kept here to document that it is never True
    network_default: str = "deny"     # the whole posture in one field

    def validate(self) -> None:
        # Fail loudly if someone "temporarily" tries to loosen the core invariant.
        assert self.mount_secrets is False, "secrets must never be mounted into a sandbox"
        assert self.network_default == "deny", "egress must be default-deny"

The validate() call is not decoration. The way these systems fail is that someone flips mount_secrets to True on a Friday to unblock one feature and nobody flips it back. Making the invariant an assertion means the unsafe configuration cannot start.

Give code a resource, never a key

The obvious objection is practical. If the sandbox has no secrets and no network, how does the agent's code read the database or call the API it was asked to use? Through a broker you own.

The pattern is a proxy that sits between the sandbox and everything else. The sandbox's code holds one address, a loopback endpoint, and nothing more. When it needs a resource, it asks the broker. The broker checks the request against the allowlist, attaches the real credential on the way out, performs the call, and logs it. The credential never enters the sandbox, so code that is fully hostile still cannot read a key it was never given or reach a host the broker will not route to.

# Runs OUTSIDE the sandbox. The sandbox can only reach it on loopback,
# and it is the only route to any real resource.
import logging
from urllib.parse import urlparse

logger = logging.getLogger("egress_broker")

class EgressBroker:
    def __init__(self, allowlist: set[str], secrets: dict[str, str]):
        self._allow = allowlist          # e.g. {"api.internal", "data.warehouse"}
        self._secrets = secrets          # real tokens, held only out here

    def fetch(self, task_id: str, host: str, path: str) -> dict:
        # 1. Enforce the allowlist. An unknown host is refused, not routed.
        if host not in self._allow:
            logger.warning("task=%s blocked egress to %s", task_id, host)
            raise PermissionError(f"egress to {host} is not allowed")

        # 2. Attach the credential here, where the sandbox can never see it.
        token = self._secrets.get(host)
        headers = {"Authorization": f"Bearer {token}"} if token else {}

        # 3. Make the real call on the sandbox's behalf and record it.
        logger.info("task=%s egress host=%s path=%s", task_id, host, path)
        return _perform_request(host, path, headers)   # your real HTTP client

Two things fall out of this for free. You get an audit line for every outbound call the agent's code made, which is the record you will want the first time someone asks what a run actually touched. And you get a real enforcement point for data leaving, which is the output side of the same problem I wrote about in agent output guardrails. Input defenses keep bad instructions out; the broker is where you decide what is allowed to leave.

Wire it into the agent as a disposable tool

Now assemble the pieces into the tool the agent actually calls. The lifecycle is the point: create a fresh box, run exactly one piece of code under the policy, capture the result, and destroy the box no matter how the run ended.

import uuid

def execute_code_tool(code: str, provider, broker: EgressBroker) -> dict:
    """The function the agent calls when it wants to run code.
    `provider` is any microVM/sandbox backend (E2B, Modal, gVisor runner...)."""
    policy = SandboxPolicy(
        egress_allowlist=("api.internal",),
        wall_clock_seconds=20,
    )
    policy.validate()                       # unsafe config cannot even start
    task_id = uuid.uuid4().hex

    box = provider.create(policy, broker_endpoint="http://127.0.0.1:8899")
    try:
        # One task, one box. The code gets loopback-only networking and no env.
        result = box.run(code, timeout=policy.wall_clock_seconds)
        return {
            "task_id": task_id,
            "stdout": result.stdout[:10_000],   # cap what comes back, too
            "exit_code": result.exit_code,
            "timed_out": result.timed_out,
        }
    finally:
        # Always destroy. No reuse, no leftover state for the next task.
        box.destroy()

The finally is doing real work. A box that survives a crash is a box the next task inherits, and inherited state is how a poisoned run reaches past its own turn. Destroy it on the error path, the timeout path, and the happy path alike.

This is also where the code-execution interface connects to the rest of your agent design. Running code as a tool is a powerful way to cut token costs, which is the argument I made in giving agents a code interpreter instead of dozens of tool schemas. That efficiency is real, and it is also exactly why the attack surface shows up: the more you lean on generated code, the more it matters that the code runs somewhere it can do no harm. And since the code is so often written from untrusted input, this sits right next to defending tool-using agents against prompt injection. Injection is the way in; the sandbox is what contains the blast when it lands.

Tradeoffs and the places this bites

None of this is free, so be honest about the costs. A microVM has a cold-start, typically tens to a few hundred milliseconds, and creating a box per task adds latency. Pre-warming a small pool helps, but a pool reintroduces the reuse risk you just designed out, so warm the box and still wipe its state between tasks rather than handing it over intact.

The egress broker is a real service you have to run and keep available, and it is on the critical path, so it needs the same care as any other dependency. The payoff is that it is also the one place you can see and control everything the generated code reaches, which is worth the operational weight.

The failure mode to watch is the slow erosion of the policy. Every one of these controls will, at some point, block a legitimate feature, and the tempting fix is always to loosen the sandbox rather than broker the new need properly. That is why the invariant lives in an assertion that refuses to start, not in a comment that asks people to be careful. Comments do not survive a deadline.

And set the resource caps against real workloads before you trust them. A twenty-second wall clock and 512MB sound generous until a genuine analysis on a large upload needs more, and now your safety limit is a reliability bug. Measure the honest tasks, set the ceiling above them, and alert when a run gets close, because a run that is about to be killed and a run that is trying to mine are easy to tell apart in advance and impossible to tell apart after.

The takeaway

A code-writing agent is a program that runs code written by whoever last influenced its input. Say that sentence out loud and the in-process exec and the plain shared-kernel container stop looking like shortcuts and start looking like the incident they are. The code is untrusted because you did not write it and cannot review it, and it is often shaped by someone hostile, so it has to run somewhere it can do nothing you did not allow.

That place has five properties and they are not optional: hard kernel-level isolation, no secrets anywhere inside, default-deny egress through a broker that holds the keys, strict CPU and memory and time limits, and a box you destroy after one task. Build those once, behind the tool your agent calls, and the worst a poisoned run can do is waste a few cents of compute in a box you were going to throw away anyway.

If your agents write and run code today, it is worth looking at exactly where that code executes and what it can reach when it does, before an uploaded file shows you the hard way. Book a consultation call and we can map your code-execution paths, find what the generated code can currently touch, and put a sandbox around it that holds.

Viral Ruparel

Generative AI consultant helping teams ship reliable LLM and agent systems in production.

Contact Viral about your AI project →

Frequently Asked Questions

Why is running AI-agent-generated code more dangerous than running my own code?+

Because you did not write it and you cannot review it before it runs. A code-writing agent generates a fresh program on every turn, often shaped by inputs you do not control: a document it was asked to summarize, a web page it fetched, a tool result from an upstream service. If any of that carries a hostile instruction, the code the agent produces can carry it out. Your own code is a fixed artifact you test once; agent code is a new, unreviewed artifact every single time, so it has to be treated as untrusted by default.

Isn't a Docker container enough to sandbox agent code?+

Usually not on its own. A standard container shares the host kernel, so every syscall the agent's code makes hits the same kernel that runs your other workloads, and one kernel escape CVE reaches the host. Containers also do nothing by default about outbound network or mounted secrets, which are the two things that actually hurt you. For untrusted code you want kernel-level isolation (a microVM or a userspace kernel such as gVisor), plus default-deny egress, no credentials inside the box, strict resource limits, and a fresh box per task. A plain container gives you none of those guarantees for free.

How does agent code reach a database or API if it has no credentials?+

Through a broker you control. The sandbox gets no keys and no open network. When the agent's code needs a resource, it calls an egress proxy that sits between the sandbox and the outside world. The proxy checks the request against an allowlist, attaches the real credential on the way out, and logs what happened. The agent's code only ever holds a loopback address, so even if that code is fully hostile it cannot read a secret it was never given or reach a host the proxy will not route to.