AI Agent Security: A Guide to Containing What Agents Can Do
Thirteen essays on AI agent security in reading order: trust boundaries, per-call authorization, scoped identity, egress checks, sandboxes, supply chain gates, and audit records.
Most of the agent security incidents I have seen did not involve a clever exploit. They involved an agent doing exactly what it was built to do, with inputs somebody else controlled. A support agent reads an email that tells it to forward the admin contact list. A data analysis agent runs code written from a spreadsheet cell that asked it to print the environment. A refund tool fires with the whole application's credentials on behalf of a user who owned none of the orders involved. Nothing crashed. Nothing threw. The agent read something, decided something, and acted, and the action was the breach.
That is the shape of the problem, and it is why the usual tools point the wrong way. Content filters inspect the words a model produces. Input sanitizers inspect what goes in. Neither says anything about the refund the agent just issued or the host its generated code just connected to. An agent is dangerous because it combines two things that are harmless alone: it reads content it did not write, and it can take actions through tools. Everything in this guide is about keeping those two things from meeting without a check in between.
I have written thirteen essays on pieces of this, each deep on one control. This guide is the calm version that puts them in order, states each pattern in plain terms, and links the essay that builds it.
Why reading plus acting is the whole attack surface
An LLM has no reliable notion of who said what. Everything arrives as tokens in one context window, and a convincing instruction in any of them gets treated with roughly the same seriousness as an instruction from you. That is not a bug you prompt away. It is how the models work, and the consensus through 2026 is that injection is not solved at the model layer and probably will not be.
For an agent that only summarizes, this barely matters. The worst an injection can do is make the summary wrong. The exposure begins when the agent can act. The moment an email tool, a payment API, or a database write is within reach of untrusted content, anyone who can place text where your agent will read it can borrow your agent's permissions. So the working strategy is containment, not prevention. Assume some injections will land and make sure a landed injection cannot do much.
Three moves do most of the work. Keep untrusted content away from the model that holds the tools: the dual LLM pattern has a quarantined, tool-less model read the hostile text and return only typed fields, so the attacker's sentence cannot survive the trip to the planner. Give every tool the least privilege it can do its job with, because an agent that cannot call the payment tool cannot be tricked into a fraudulent payment. And put a deterministic policy layer, plain code that cannot be argued with, in front of any action that is hard to undo. I went through all three, and why fencing in the system prompt is a hint rather than a wall, in prompt injection defense for tool-using agents.
One principle from that essay runs through everything that follows: if a control lives only in the prompt, assume it can be bypassed, and put the control that matters in code.
Authorize every action, not every tool
The second half of the problem shows up even without an attacker. An agent that holds a tool runs it with the application's credentials, not the current user's. A support agent with an issue_refund tool wired to a service account can refund any order in the system, and nobody checks whether this user, in this conversation, was allowed to refund that order. This is the confused deputy: a privileged process tricked by a less privileged party, which can be a malicious document or an honest customer asking for something they are not entitled to.
The trap is conflating two questions. "Can the agent invoke this tool" is about wiring: the function is registered, the schema is valid. "Is the current user permitted to perform this action on this resource" is authorization. Most frameworks answer the first and silently treat it as the second, so registering a tool becomes granting permission to everyone the agent ever talks to.
Closing the gap takes three moves. Carry the authenticated principal from the session into every tool call, captured from trusted scope and never passed as an argument the model fills. If the model can set user_id, an injection will set it to whoever it likes. Put one fail-closed gate in front of every side-effectful action that loads the real resource and checks the user's real grants against it, defaulting to deny when no rule matches. And hand each tool a credential scoped to just that call, so a bug in the gate is still contained by a key that only works for one customer for sixty seconds. Authority is the intersection of what the tool can do and what the user may do, computed from grants, not from the conversation. The full treatment is in the confused deputy essay.
Put that gate at the one chokepoint every tool call already passes through, the harness code that invokes the function, not inside the tools, where the first tool a teammate adds without the check becomes the hole. The decision maps an action context, built by the harness from server-side state, to allow, deny, or escalate to a person.
type Decision =
| { effect: "allow" }
| { effect: "deny"; reason: string }
| { effect: "escalate"; reason: string; approvers: string[] };
type Rule = (ctx: ActionContext) => Decision | null; // null = abstain
const rules: Rule[] = [
(ctx) => ctx.tool === "delete_account"
? { effect: "deny", reason: "no agent may delete accounts" } : null,
(ctx) => {
if (ctx.tool !== "issue_refund") return null;
const cents = Number(ctx.args.amountCents ?? 0);
if (!Number.isFinite(cents) || cents < 0)
return { effect: "deny", reason: "refund amount is not a valid number" };
if (cents > 50_000)
return { effect: "escalate", reason: "over the auto-approve limit", approvers: ["finance-oncall"] };
return { effect: "allow" };
},
];
export function decide(ctx: ActionContext): Decision {
for (const rule of rules) { const d = rule(ctx); if (d) return d; }
return { effect: "deny", reason: "no policy permits this action" };
}
The last line is the design. Default deny means a tool someone ships next month is refused until a human writes a rule for it. Wrap the evaluation so a crash in the policy code also denies, log every decision including the allows, return a deny to the model as a structured tool result rather than an exception, and run in shadow mode first so you can read the would-have-denied rate before enforcing. The wiring and the rollout are in the policy-as-code gate.
Give the agent an identity before you trust its calls
Every gate above assumes it knows who is calling. Most agents cannot say. A recent survey put the number at 93 percent of agent projects authenticating with a single unscoped API key. It works, which is the trap, because authentication fails silently until it fails all at once.
The costs show up the day something goes wrong. Every call the agent ever made looks identical in downstream logs, so "for whom" is unanswerable. The key holds the union of every permission any user might need, so one leak or one steered action exposes everything the agent can do. And an authorization policy reasoning about a caller that is always the same anonymous key is reasoning about a fiction.
The fix is to treat the agent as what it is, a non-human identity acting on behalf of humans, and to carry two facts with every call: which software is calling, and on whose behalf. The agent gets a workload identity, a SPIFFE SVID rather than a static secret in an environment variable. OAuth 2.0 token exchange, RFC 8693, combines them. The agent trades the user's token as subject and its own token as actor for a short-lived token whose subject is the user, whose nested act claim names the agent, whose audience pins it to one downstream service, and whose scope is the least privilege for this one action. The resource server then authorizes and logs against both parties, and an incident review finally reads "agent X acted for user Y with scope Z."
Put the exchange behind the single function every outbound tool call must use, and make the raw shared key unreachable from tool code, or the scheme is theater. Short-lived tokens will expire mid-task, so re-exchange on demand and keep side effects idempotent. If the full SPIFFE stack is more than you can take on today, a static token per environment scoped to one service at least kills the god key. The exchange, the cache, and the resource-side check are in giving your agent a scoped identity instead of a shared key.
Check what leaves, not just what enters
Everything so far faces inward. The one artifact an agent produces that reaches a customer, an inbox, or a public channel is its output, and it is the one artifact you cannot take back. A harmful reply does not need a hostile input. The model retrieves a record it was allowed to retrieve and surfaces a field it should not have. Or it invents a refund policy that sounds plausible and commits you to it in writing. The instructions were clean; the output was not.
The fix is a fail-closed egress layer: a single stage between the model's finished output and the send, running an ordered list of checks. Each check can pass the text, rewrite it by redacting a span, or block it outright. Cheap deterministic checks run first: card-shaped numbers confirmed with Luhn, email addresses that are not the recipient's own, links outside an allowlist. A model-based judgment for the fuzzy cases, like a commitment nobody authorized, runs last and only on text that survived. Any exception blocks.
def run_egress(text: str, ctx: dict, checks: list[Check]) -> CheckResult:
current = text
for check in checks:
try:
result = check(current, ctx)
except Exception as e:
# A check that throws must not open the gate. Fail closed.
return CheckResult(Action.BLOCK, current, f"{check.name} errored: {e}")
if result.action is Action.BLOCK:
return result # stop immediately, nothing ships
current = result.text # carry redactions forward
return CheckResult(Action.ALLOW, current)
When a check blocks, the caller sends a safe fallback and logs the draft for review, or routes it to a person to release. Two traps are worth naming. Streaming raw tokens defeats the guarantee, because you have shipped the first eight digits of a card before the check sees the ninth. Buffer to a sentence boundary, check the unit, release it. And do not log the thing you just redacted, or you recreate the leak in a system that is usually less protected than the one you were guarding. The checks, the streaming buffer, and the precision problem of over-redaction are in agent output guardrails.
Run generated code somewhere it can do no harm
A code-writing agent is a program that runs code written by whoever last influenced its input. Most teams run that code in-process or in a plain container, where os.environ is your secrets, open() is your filesystem, and socket is your network. The code does not need to be malicious to walk through those doors. It needs to be convinced that walking through is part of the task.
A sandbox worth the name enforces five things, and skipping any one is the gap an incident walks through. Hard, kernel-level isolation: a microVM or a userspace kernel like gVisor, not a container sharing the host kernel. No credentials inside the box, of any kind, scoped or not. Default-deny egress, so stolen data has no way to leave. Hard caps on CPU, memory, and wall-clock time. And a short life: one task, one box, destroyed in a finally regardless of how the run ended, because inherited state is how one user's poisoned run reaches the next user's.
The practical objection is how code reaches a database with no keys and no network. Through a broker you own. The sandbox holds one loopback address and nothing else. The broker checks each request against an allowlist, attaches the real credential on the way out, performs the call, and logs it. The credential never enters the box. Put the invariants in an assertion that refuses to start rather than a comment asking people to be careful, because comments do not survive a deadline. The policy object, the broker, and the disposable tool wrapper are in sandboxing agent code execution.
Treat every capability you did not write as untrusted
Agents grow by pulling in things from outside: skills off a registry, tools from an MCP server, calls to another team's agent over A2A. Each one extends your trust boundary to code and identity you do not control, and each one has the same fix, which is an admission gate that nothing reaches the agent without passing.
Third-party skills come first because the gate is cheapest. A manifest that says "read-only" is a text file the author typed. A version tag is a label the registry can move. So pin the content hash of the whole skill directory and refuse to load anything that does not match the pin a human recorded at review, which turns a silent swap into a diff someone signs off on. Then scan the code for calls that reach outside the process and compare them against what the manifest declares. The scan is a tripwire, not a proof, and it catches the common case of a useful skill with one extra socket bolted on. And run the skill in a box that can only reach the hosts it declared, with the real egress control at the OS and proxy layer where the skill cannot patch it back. All three gates are in verifying and sandboxing third-party skills.
MCP tools have a sharper problem: you approved them once, and the spec lets the server return a different tools/list on the next connection with no re-approval and no integrity check. The tool you vetted on Monday can ship a poisoned description on Friday. The fix is a lockfile for tool definitions. At approval, fingerprint the three fields the model reads to decide behavior, canonicalized so key order and whitespace do not fire false alarms.
import hashlib, json
def fingerprint_tool(tool: dict) -> str:
canonical = {
"name": tool.get("name", ""),
"description": tool.get("description", ""),
# sort_keys makes a reordered schema hash identically,
# so only a real change in meaning reads as drift.
"input_schema": tool.get("inputSchema", {}),
}
blob = json.dumps(canonical, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(blob.encode("utf-8")).hexdigest()
Then refetch the live definition every session, compare before every call, deny any tool not in the approved set (a rug pull can add a tool as easily as mutate one), and route drift to a human with the old and new descriptions side by side rather than crashing the run. Fingerprinting cannot catch a server that was malicious from the first fetch, so it belongs behind a server allowlist and in front of the default-deny action gate. That is the MCP rug pull essay.
Other agents are the widest boundary. An Agent Card fetched over perfect HTTPS tells you the bytes came from a hostname, not that the host may speak for the agent it names or that the endpoint inside the card belongs to the real provider. Verify the card's JWS signature against an issuer key you pinned out of band, never from the host that served the card. Pin the hash of the exact card you reviewed, so a rotated endpoint or a weakened auth scheme stops the call instead of silently redirecting it. Do not forward your own access token; mint a fresh, audience-bound, short-lived one per call through the same token exchange described above. Bind each request to a nonce and a timestamp so a captured call cannot be replayed into a duplicate charge. And treat the remote agent's response as untrusted input, because a verified agent can still return a poisoned instruction. The admission gate is in A2A agent card verification.
Put a person in front of the one-way doors
Some actions are not worth automating fully no matter how good the model is: the refund, the production deploy, the email to a customer, the delete. Average-case accuracy does not make a one-way door safe to walk through blind. The honest move for that small set is to let the agent do all the reasoning, then stop it just before the irreversible step and ask.
The failure mode of human-in-the-loop is not too little oversight but too much. A system that prompts before every tool call trains reviewers to click approve on reflex, and a rubber stamp is worse than no gate because it manufactures a false record of human judgment. So the gated list is short and the rule is reversibility: if the agent gets this wrong, can you quietly undo it before it matters? Reads and drafts run free. Money, production changes, outbound messages, deletes, and anything legally binding stop.
The mechanic is propose, pause, resume. The agent emits the exact arguments and a rationale, the run persists a pending record and releases the worker, and when a decision arrives the run is rehydrated from its checkpoint and continued. The human approves the exact args, so what executes is byte for byte what a person saw. Resolution is one atomic state transition, first writer wins, so a double click or two reviewers acting at once execute the side effect once. A pending record nobody answers times out to reject and escalate, never to approve. And the agent never waives its own gate on a confidence score. This is where the escalate branch of the policy gate lands, and the buildable version is the approval gate essay.
Keep records that answer why and who
A trace with forty-seven spans shows that something ran. It does not show what the agent decided, what rule it applied, what facts it acted on, or who stands behind the action. A trace is telemetry for the engineer who already knows the system. An audit trail is for someone who was not in the room and shows up months later with authority over whether you keep operating. You cannot reconstruct intent from a latency waterfall, and for high-risk systems under the EU AI Act, the Article 12 record-keeping and Article 14 human oversight obligations are being enforced now rather than discussed.
The unit you audit is the consequential decision, not the span. Emit a decision record at the moment the agent commits to the action, from the same chokepoint where the policy gate already decides allow, deny, or escalate, because the decision and the record are one event seen from two angles. The record holds a stable action id, the agent and the scoped authority it acted under, the inputs the decision rested on, the rationale captured as a required field of the structured call rather than solicited afterward, the policy decision and the rule version, the human who approved it if one did, the outcome, and the hash of the previous record so the chain is tamper-evident. Write it to an append-only store. Do not try to build this by scraping traces after the fact, because the rationale was never captured and a scraper cannot recover it. That is the audit trail essay.
Two more records round out the governance layer. The first is the inventory. A year of shipping leaves most teams with a fleet nobody can list: no owner, no declared budget, no record of which tools each agent reaches. The fix is a registry. A small typed manifest every agent must declare, with an id, a routable owner email, a purpose, the model, the permitted tools, a daily budget, and a lifecycle state. A CI gate that refuses to deploy anything unregistered, retired, or malformed. A runtime that reads the budget and fails closed. And a reconciliation job that compares declared tools against the tools the agent was observed calling. Issuing workload identities assumes you know which agents exist, so the registry comes before the identity work. It is in ending agent sprawl with a registry.
The second is the mark on the output itself. Since August 2, Article 50 of the EU AI Act requires that synthetic images, audio, video, and text carry a machine-readable marking that they were made by AI. A caption in your UI does not travel with the JPEG when a user saves it. The settled answer is a C2PA Content Credential: a signed manifest embedded in the file with digitalSourceType set to trainedAlgorithmicMedia, applied at the egress boundary as the single return path from your generation service, verified in an integration test so the obligation cannot regress, and paired with a stored fingerprint so a stripped file can be matched back. The build is in marking AI output for Article 50.
Where to start
If you have none of this, do not try to build all of it at once. The order below is the one I would use, and each step makes the next one meaningful.
-
Write down what each agent can touch. Before any control, produce the list: which agents exist, who owns each, what tools each holds, where untrusted content enters. The registry manifest is the cheapest form of this, and the CI gate keeps it from going stale.
-
Scope the tools. Least privilege is the highest-value control and the most boring. Remove tools the agent does not need. Replace "send email to any address" with "reply to the existing thread." An agent that cannot call the payment tool cannot be tricked into a payment.
-
Put one gate at the tool boundary. Route every tool call through a single function that builds the context from the session, evaluates policy as code, defaults to deny, fails closed, and logs every decision. Ship it in shadow mode, read the rates, then enforce.
-
Gate the one-way doors with a person. Pick the short list of irreversible actions, make the gate's escalate branch persist a pending record, and make resumption idempotent. Keep the list short enough that approvals get read.
-
Give the agent an identity. Kill the shared key. If the full token-exchange stack is too much today, start with a static token per environment scoped to one service, and plan the delegated version.
-
Add the egress check. A fail-closed output layer with deterministic redaction first and model judgment last. If you generate media for EU users, the C2PA signing step goes in the same place.
-
Contain the code and the capabilities. Move generated code into a real sandbox behind a broker. Pin and scan skills. Fingerprint MCP tools and refetch every session. Verify agent cards and mint per-call tokens for A2A.
-
Emit decision records. Once the gate exists, the audit record falls out of the same seam. Add it before an auditor asks, because retrofitting accountability onto a year of raw spans is far more expensive.
None of this assumes you can stop an injection from landing, a skill from lying, or a model from summarizing the wrong row. It assumes you cannot, and it makes sure that when one of those happens it hits a wall instead of your production data. Every control here is a few hundred lines and a habit of routing things through one door. The payoff is that the gap between an agent that demos well and one you can put in front of the internet stops being a matter of luck.