Context engineering for AI agents: the production guide

One reference for the whole discipline: what goes into an agent's context window, what gets pushed out, and the patterns that keep long runs cheap, true, and auditable.

Guide12 essays20 min readUpdated

Most of the agent failures I get asked to look at are not model failures. The model is fine. What it was handed was wrong: a transcript that grew for three hours until the plan from turn three fell out of it, a tool result from twenty turns back describing a file that has since changed twice, a subagent summary that dropped the one warning that mattered, a memory store that confidently recalls a decision the user reversed in March. In every one of those cases the fix was not in the model. It was in what went into the context window and what was allowed to stay there.

That is context engineering: deciding, in code, what the model sees on each call. What enters, what gets compressed, what gets evicted, what lives outside the window and is fetched on demand, and what carries over to the next session. The term became fashionable this year as context windows crossed a million tokens, which is backwards, because a bigger window does not make the problem go away. It makes the failures quieter. A request that used to error out now returns a confident answer that ignored what you buried at forty percent depth.

I have written about each piece of this separately, usually after watching the same break happen in more than one system. This guide is the calm version: the parts in order, the pattern that fixes each one, and a link to the essay that goes deep.

The window is a budget, not a bucket

Start with the number everyone gets wrong. The advertised context window is the number of tokens the API will accept without an error. The effective context window is the length past which the model stops using the input reliably: recall of facts in the middle drops, instructions get skimmed, answers contradict the source. The first is a hard limit. The second is a quality cliff you have to find yourself, and a model that accepts a million tokens often holds clean recall to around half to two thirds of that.

Two curves cross here. Cost is linear: in an agent loop every turn resends the accumulated context, so a 200K prompt is 200K on this turn and the next and the one after. Quality is not linear. It degrades as input grows, worst in the middle. Past some length every token you add costs money and lowers the odds the model uses the tokens that matter.

The pattern is to measure, then budget. Run a positional recall probe against the model you actually ship: plant a fact the model cannot know from pretraining at a known depth in a realistic document, pad to a target length, ask for it back, and sweep both depth and length. Where recall falls below your bar is your effective limit. Then assemble every prompt as a budgeting problem below that line.

from dataclasses import dataclass

@dataclass
class Chunk:
    text: str
    tokens: int
    priority: int   # higher wins; the system prompt and the question are top

def assemble(chunks: list[Chunk], budget_tokens: int) -> list[Chunk]:
    # Greedy fill by priority. What does not fit is dropped and reported.
    ordered = sorted(chunks, key=lambda c: c.priority, reverse=True)
    kept, used = [], 0
    for c in ordered:
        if used + c.tokens <= budget_tokens:
            kept.append(c)
            used += c.tokens
    dropped = len(chunks) - len(kept)
    print(f"context: {used}/{budget_tokens} tokens, {dropped} dropped")
    return kept

The budget you pass in is not the advertised window and not even your measured ceiling. It is the ceiling minus room for output and a safety margin, because the measured limit is where quality starts to slip, not where it collapses. If recall held to 400K, budget 250K to 300K and treat the rest as headroom. Put a guard where the prompt is finalized that emits utilization into your traces, so overflow is a metric you alert on instead of a support ticket you cannot reproduce. The probe, the budget, and the guard are in the 1M token window trap.

Everything that follows is a way of spending that budget well.

Keep the history from growing without bound

The number one way a long-running agent breaks is that its history grows until it either hits the ceiling or starts hallucinating decisions it made an hour ago. The lazy fix is truncation, and dropping the oldest messages is the fastest way to make the agent forget the plan it committed to on turn three.

The pattern that holds up is anchored summarization. Keep one persistent anchor summary of everything old. Keep the last four to eight turns exactly as they happened, so recent detail is never lossy. When the total crosses a token threshold, fold the turns about to age out into the anchor. Merge them in; do not rebuild the summary from scratch, because re-summarizing the entire history on every pass lets the summary drift the way the raw transcript would.

A few rules separate a compaction loop that works from one that quietly wrecks things. Measure with the provider's token-counting endpoint, not a client-side guess. Compact only once you cross the threshold, never on a timer. Never compact between a tool_use block and its tool_result. Tell the summarizer what must always survive: decisions, IDs, file paths, open TODOs. And append the anchor to the end of the stable system prompt rather than splicing it at the front, so the prompt cache prefix survives across turns.

If you would rather not babysit the summarizer, the Anthropic API has server-side compaction in beta. The integration bug that catches almost everyone is appending only the text of the response back to the message list. The response carries compaction blocks the API needs on the next request, so append resp.content whole or the feature silently does nothing. The full loop, the managed variant, and the pitfalls are in context compaction for long-running agents.

Keep the context true, not just small

Compaction answers "how do I make this context smaller." A different question is "how do I make this context true," and conflating them is why teams reach for the wrong fix.

An agent reads a config file on turn six, edits it on turn nineteen, edits it again on turn thirty-one. On turn forty it needs the current contents. It holds three versions, all equally authoritative, and reaches for the one from turn six. Nothing errored. The token count is fine. The agent acted on a world that stopped existing twenty-five turns ago. This is agent drift, and it is a correctness bug. In one evaluation of long-horizon tool-using agents, roughly a fifth of tasks hit the turn limit with the context still around seven thousand tokens. They were not running out of window. They were confused inside one full of contradictory snapshots.

You can summarize that transcript perfectly and still carry the lie forward, because the fact was true when recorded and the summarizer has no idea it was contradicted. Shrinking a context that is lying to you gives you a smaller, cheaper lie. The lost-in-the-middle effect makes it worse: the stale observation sits in the weakly-attended middle, exactly where the model grabs without scrutiny.

The pattern is to stop treating tool results as an append-only log and start treating them as claims about resources, each with exactly one current claim. Key every observation by what it describes: a file by its path, a record by table plus id, a fetch by URL. A newer observation of the same resource supersedes the older one, which moves to an audit list the model never sees. Render the context from the live set only. Observations that are old but not superseded stay, flagged so the model re-reads before acting. That is pruning for what was replaced and staleness marking for what merely aged.

Two things you must not do. Do not collapse events the way you collapse snapshots: send_email is not a snapshot of anything, and two sends to the same address are two distinct events. Collapse them and the agent forgets it already did something, then does it again. And do not prune the agent's own reasoning, plan, or record of actions taken; dropping those is amnesia, which is worse. The store, the key function, and the render step are in the agent is not confused, its context is stale.

Stop large payloads at the tool boundary

A tool call returns 38,000 tokens of account JSON. The agent needed one boolean in field fourteen. That would be wasteful once, and it is not once: the blob is now in the history and every turn after it re-sends all 38,000 tokens. Three big responses early in a ten-step run means paying for all three on all ten steps.

Compaction will summarize that blob eventually. The cheaper move is to never let it in. When a program opens a large file it gets a descriptor, not the whole file by value. Do the same at the tool boundary: wrap every tool, and if a result is over a length threshold write the payload to a store and return a handle plus a preview.

import json, uuid
from pathlib import Path

STORE = Path("/var/agent/artifacts")
INLINE_LIMIT = 2000   # chars, roughly 500 tokens

def offload_if_large(tool_name: str, raw_output: str) -> dict:
    if len(raw_output) <= INLINE_LIMIT:
        return {"inline": True, "content": raw_output}
    artifact_id = f"{tool_name}-{uuid.uuid4().hex[:8]}"
    STORE.mkdir(parents=True, exist_ok=True)
    (STORE / f"{artifact_id}.txt").write_text(raw_output)
    return {
        "inline": False,
        "artifact_id": artifact_id,
        "approx_tokens": len(raw_output) // 4,
        "preview": build_preview(raw_output),   # shape, not values
    }

The preview is the part that matters. For structured data it exposes the shape: top-level keys, array lengths, a couple of sample rows. For text and logs it is the first and last lines plus a total count. It is a table of contents, not a teaser; a thin preview is worse than no offloading, because the model either fetches the whole thing back or guesses from incomplete data. Pair it with a read_artifact tool that takes the handle and a slice address, a dotted key path for JSON or a line range for text.

Watch for two things in your traces. If the model fetches the same artifact on every turn, it should have been inline. And handles go stale the same way observations do, so key artifacts by resource and let a newer write supersede the older handle. Offloading and compaction stack: offload the blobs so they never bloat the history, then compact the history that remains. The wrapper, the preview builder, and the read-back tool are in your tool returned 40,000 tokens, the agent needed 12.

Retrieve what earns a place in the window

A bigger window changes the economics of retrieval. It does not remove the reason for it. A few thousand relevant tokens still beat a hundred-thousand-token dump on the same task, at a fraction of the cost. So what goes into the window is largely a retrieval question, and it applies to two things most teams think of as one.

The first is documents. When a RAG system gives a wrong answer, the model is usually not the problem; the right chunk never made it into the prompt. Dense embeddings are good at meaning and bad at specifics: error codes and function names carry almost no semantic weight, so the one page that names ERR_CONN_4021 does not stand out from twenty pages about connection errors. Keyword search has the mirror problem. Run both, each to a real depth, and merge with reciprocal rank fusion, which needs no calibration. Then rerank the fused candidates with a cross-encoder, with a threshold so a question your corpus cannot answer returns nothing instead of three loosely related chunks the model will improvise from. Measure recall at k on real questions from your own traffic, stage by stage. The code and the latency numbers are in your RAG is not broken, your retrieval is.

The second is tools. Connect enough MCP servers and the agent carries two hundred tool definitions into every turn. It pays for all of them on every request and still picks the wrong one, because choosing among forty near-duplicate schemas is the lost-in-the-middle problem pointed at your tool list. Treat tools the way you treat documents: embed each description once at startup and retrieve the handful that match the task. For multi-turn runs that is a small working set maintained across the run, seeded with always-on core tools, with recently used tools kept sticky and the rest evicted. Add one meta-tool, always in the core set, that lets the model search for a capability by intent when nothing visible fits. And write tool descriptions for retrieval, in the user's language; that is the highest-value change and the one most teams skip. The index, the working set, and the escape hatch are in your agent has 200 tools and picks the wrong one.

Draw the boundaries between agents on purpose

With more than one agent, context engineering becomes a question of what crosses between windows. Two patterns cover most of it.

The first is subagent isolation. Spawning a subagent to do the messy work in its own fresh window and return only a result is the cleanest answer we have to context rot on long runs. The parent's context grows by a paragraph instead of a thousand tokens of grep output. But that paragraph is a lossy compression of everything the subagent saw, with no schema and no provenance, and the moment it returns the subagent's context is gone. When the parent decides a migration is safe and it was not, the transcript says only "the schema check reported no blocking issues," and the warning the subagent chose not to mention is unrecoverable.

The fix is not a smarter summary. It is a return contract. Make the subagent fill named fields instead of writing prose: a verdict, a one-line summary, caveats for what did not get checked, a confidence score it has to commit to, and evidence handles. The caveats field matters most, because a model answering "what did you find" reports what it found, not what it skipped, and a named field forces the omission into the open. The evidence handles are offloading applied to the agent boundary: raw findings go to a store, references come back, and the audit trail outlives the subagent. And do not spawn reflexively. Every call pays a summarization tax, so inline small work and reserve isolation for tasks whose noise is large relative to the conclusion. The contract, the store, and the spawn policy are in your subagents are hiding the evidence.

The second pattern is the one you have probably built by accident. Direct handoffs work for three agents. Add a fifth and every function signature takes half the other agents' outputs, so you introduce a shared state dict that everyone reads and writes. That works until the dict has thirty keys, three agents write to results, and a value that looks current is two rounds stale. That dict is a blackboard, a pattern from the 1970s, and it became a mess because you built a quarter of it. The full pattern has a board, knowledge sources that know when they have something to contribute, and a control component that decides who runs next. Build it deliberately: typed entries that record their author, their parents, and what they supersede; agents that never mutate an entry but write a new one pointing at the one it replaces; a live view with superseded entries hidden; and a control loop with a step budget. Scope each agent's reads to the entry kinds it needs, because pulling the whole board into every prompt is how blackboards get expensive. The rules are in your multi-agent system already has a blackboard.

Carry the right residue across sessions

Everything so far is about one run. Models are stateless, so when the session ends the context is gone and the agent greets your best user tomorrow like a stranger. A bigger window does not fix this; it still starts empty on restart. Memory is a layer you build outside the model, in three stages, and most teams stop after the first.

The first stage is persistence: extract, store, retrieve, consolidate. After a conversation, ask the model for durable, standalone facts, with explicit permission to return nothing, because an extractor obliged to produce something slowly fills the store with noise. Embed the facts and store them scoped by user; memory that leaks across users is a privacy incident, not a quality bug. At the start of a turn, pull the three or so most relevant facts and put only those in the prompt. Then consolidate on write: check a new fact against what you hold for that subject and decide whether to add, update, or ignore, so a user who moved from React to Svelte is not stored as both. The loop is in your agent forgets you the second you close the tab.

The second stage is the one people discover six months later. The store holds the user's timezone eleven times, three contradictory plans because they upgraded twice, and a pile of "user seemed frustrated" that was never precise enough to act on. Retrieval runs over all of it and quietly gets worse every week, and a stale fact often outranks the fresh one because it was recorded more times. This is a garbage-collection problem. Gate what gets written: durable, specific, confidently extracted, default no. Reconcile and dedupe on the way in, with single-valued predicates like plan superseding rather than accumulating, and superseded rows marked rather than deleted. Decay what stops earning its slot with a score built from idle time, real use, and confidence, with a pin for facts that must never go. And measure duplicate rate, contradiction rate, and retrieval hit rate, because a store where most retrieved tokens are never used is paying to poison its own prompt. The curation layer is in your agent's memory is full and most of it is junk.

The third stage is memory of how to do the work, not facts about the user. Your agent mishandles the same multi-currency invoice every Monday, you add a line to the system prompt, and three months later the prompt is 4,000 tokens of patches that contradict each other. The obvious alternative, summarize each task into a note and rewrite the blob, collapses in two ways: brevity bias drops the specific caveat that made the lesson worth keeping, and rewriting everything on every update compounds small distortions until the playbook has forgotten most of what it knew. Agentic context engineering is the structural fix. A Generator does the task, a Reflector extracts a few concrete lessons from the trajectory without seeing the existing playbook, and a Curator appends them as small deltas, reinforcing near-duplicates instead of adding them, and prunes on a schedule rather than every step. Grade each item with helpful and harmful counters from real outcomes so useful lessons rise and wrong ones retire, and treat the playbook as an attack surface, because anything that shapes future behavior can be poisoned by untrusted input. The three roles are in your agent fails the same way every week and learns nothing.

Handle reasoning traces as context too

The newest thing in the window is the model's own thinking. Turn on extended thinking and the response is an ordered list where thinking blocks precede the answer, billed as output tokens whether you display them or not. Most teams flipped the flag and changed nothing about how they read the response or write their logs, and that gap is where three failures live.

The first is cost: a classification call that was a hundred cheap tokens now spends several hundred output tokens on a decision that was never hard. The lever is effort, set per task rather than globally, and metered by task kind. The second is privacy: a trace restates the input in the model's words, so the account number the user pasted reappears inside the thinking, and a logger that serializes the full content array writes it somewhere your redaction does not cover. Strip thinking at the egress boundary and keep only a hash. The third is the subtle one.

async function agentLoop(initial: Anthropic.MessageParam[], tools: Anthropic.Tool[]) {
  const messages = [...initial];
  while (true) {
    const res = await client.messages.create({
      model: "claude-opus-5",
      max_tokens: 16000,
      thinking: { type: "adaptive" },
      tools,
      messages,
    });
    // Append the whole content array, thinking blocks included.
    messages.push({ role: "assistant", content: res.content });
    if (res.stop_reason !== "tool_use") return res;
    const results = res.content
      .filter((b): b is Anthropic.ToolUseBlock => b.type === "tool_use")
      .map((call) => ({ type: "tool_result" as const, tool_use_id: call.id,
                        content: runTool(call.name, call.input) }));
    messages.push({ role: "user", content: results });
  }
}

Within a tool-use turn on the same model, the thinking blocks are the reasoning the model has already started. Rebuild the assistant turn from the text and the tool call alone and the model loses its plan on every tool round, which looks exactly like the model getting dumber mid-conversation. Once a turn ends, old thinking is dead weight you pay to resend, and the API offers a context edit that clears it. Preserve inside the turn, prune between turns. The response grew a new kind of content, and content in the window has to be governed. The three failures and their handling are in your model started showing its work.

Where to start

A team with none of this does not need all of it at once. This is the order I would follow, because each step is cheap, measurable on its own, and makes the next easier to see.

  1. Measure your effective context and set a budget. Run the recall probe against the model you ship, write the number down, and put a guard at prompt assembly that emits utilization. Until you have this, you cannot tell whether any later change helped.

  2. Offload large tool outputs. A wrapper at the tool boundary and one read-back tool. On a tool-heavy agent it is the largest cost and window win available. Spend the effort on the preview.

  3. Add compaction. Anchored summarization with four to eight verbatim recent turns, triggered by a real token count, at turn boundaries only. Or opt into server-side compaction and append resp.content whole.

  4. Key observations by resource. If your agent ever sees the same thing twice in a run, add the observation store and render from the live set. This is usually where baffling decisions go away.

  5. Fix retrieval before touching the prompt. Hybrid search with rank fusion, a reranker with a threshold, recall at k on your own traffic. If the agent carries more than a few dozen tools, retrieve those too.

  6. Put contracts on the boundaries between agents. Typed subagent returns with caveats and evidence handles. If you have a shared state object, admit it is a blackboard and give it provenance, supersession, and a control loop.

  7. Build memory, then curate it. Extract, store, retrieve, consolidate, scoped by user with the delete path built on day one. Add admission control, decay, and the health report before the store is six months old.

  8. Let the agent keep a playbook. Reflector and Curator, delta updates, scheduled pruning, outcome-graded items. Last, because it only compounds on a context that is already budgeted, true, and small.

  9. Govern reasoning traces. Per-task effort, egress stripping, verbatim append inside tool turns, clearing between them.

The thread through all nine is the same. The context window is not a buffer that fills on its own while the model copes with whatever lands in it. It is the one input you fully control, and every token in it should be there because you decided it earned the place: measured against a budget, current rather than stale, a handle rather than a payload, retrieved rather than dumped, typed at every boundary it crosses, and curated rather than accumulated when it outlives the run. Do that and the model you already have starts behaving like the one you thought you were paying for.

Building this?I help teams put this into production. Thirty minutes, no pitch.