Your Agent Retries Are Making the Failure Worse
When an agent step fails and you retry it in the same context, the failed attempt stays in the window and the model conditions its next try on its own mistake. The retry does not get a fresh shot, it gets a biased one, and it often repeats or compounds the error while you pay for every token. This is how context contamination kills retry recovery, and how to fix it with clean context forks, distilled feedback, and a retry policy that knows when to stop.
A team I was helping had an agent that failed about one task in twelve, which they were fine with. What they were not fine with was that their retries almost never recovered. The orchestrator caught the error, retried the same step up to three times, and the second and third attempts failed at roughly the same rate as the first. They were paying for three attempts and getting the reliability of barely more than one. The obvious read was that the model just was not good enough at these tasks. That was not it.
The problem was in how they retried. When a step failed, they appended the error to the running message list and asked the model to try again. That feels like the right thing to do. Show the model what went wrong, let it correct. But the failed attempt was still sitting in the context, so every retry was conditioned on the mistake it was supposed to fix. The model was not getting a fresh shot at the task, it was getting a shot that had been primed toward the exact approach that had already failed.
This is context contamination, and it is one of the quieter reliability killers in production agents because it hides inside code that looks completely reasonable.
Why a contaminated context cannot recover
A language model generates each token from everything in its window. It does not have a switch that says "ignore the part above where you got it wrong." The failed reasoning, the broken tool call, the stack trace, all of it becomes part of the prior the model samples from on the next attempt. So when you leave the failure in place and append "that failed, try again," the most probable continuation is something that looks a lot like what it just did. You asked for a correction and the token statistics are pointing back at the mistake.
There is a recent paper that puts a formal frame around this, showing with an information-theoretic argument that retries in a contaminated context have worse convergence than fresh attempts, because the model has to simultaneously hold what it tried, why that was wrong, and what to do instead, and that divided conditioning undermines the recovery. You do not need the math to feel it in production. You see it as a retry that produces a near-duplicate of the failed call, or one that over-corrects into a different kind of wrong because it is trying to get away from an error that is still staring at it.
There is a cost dimension too, and it is not small. Every contaminated token rides along on every subsequent attempt. A step that failed with a long tool output now carries that output through three retries, so your "cheap" retry is actually your most expensive call of the task. And if the failed trace is long enough, it can push the live reasoning toward the end of the window where attention is thinnest, which is its own reliability tax on top of the priming.
The pattern that causes it
Here is the shape of the code that creates the problem. It is the natural thing to write, which is exactly why it is everywhere.
# The contaminated retry: every attempt sees the previous failure.
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": task},
]
for attempt in range(3):
response = model.run(messages, tools=TOOLS)
result = execute(response.tool_call)
if result.ok:
break
# This is the mistake. The failed call and its error stay in `messages`,
# so attempt N+1 conditions on every prior failure in the window.
messages.append({"role": "assistant", "content": response.raw})
messages.append({"role": "tool", "content": f"ERROR: {result.error}"})
By the third pass through the loop the model is reading its own two previous failures before it writes the third one. Each attempt is strictly more primed toward the failing region of the output space than the last. The loop is not giving the model three chances, it is giving it one chance followed by two increasingly biased echoes.
Retry on a clean context, not the poisoned one
The fix is to stop reusing the window that failed. Each retry gets a fresh context built from the original task plus a short, distilled note about what went wrong. The raw failed transcript never makes it into the window the next attempt reasons from. It goes to your logs, where it belongs, not into the prompt.
def run_step(task: str, feedback: str | None = None) -> StepResult:
# Build a clean window every call. No prior attempt lives here.
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": task},
]
if feedback:
# One compact correction, not the whole broken transcript.
messages.append({
"role": "user",
"content": f"A previous attempt failed. Avoid this: {feedback}",
})
response = model.run(messages, tools=TOOLS)
result = execute(response.tool_call)
return StepResult(ok=result.ok, response=response, error=result.error)
The important move is that feedback is a short string, not a message history. The model learns the one thing it needs to do differently and gets to approach the task without rereading the approach that failed. In practice this is the single change that turns a retry rate of "about the same as the first try" into retries that actually recover, and it does it while cutting tokens rather than adding them.
Distill the failure into feedback, not a transcript
The quality of a clean retry lives in that feedback string. Dumping the raw error back in is better than the full contaminated window but still worse than a real distillation, because raw errors are noisy and often restate the failed approach. You want the smallest accurate instruction that would change the next attempt.
You can distill with a cheap, fast model on a tight prompt. This is a good place for a small model because the job is narrow: read a failure, write one corrective sentence.
DISTILL_PROMPT = """You are debugging one failed agent step.
Given the task, the attempted action, and the error, write ONE short
instruction (max 25 words) telling the next attempt what to do differently.
State the correction only. Do not restate the failed approach or explain."""
def distill_feedback(task: str, response, error: str) -> str:
out = small_model.run([
{"role": "system", "content": DISTILL_PROMPT},
{"role": "user", "content": (
f"Task: {task}\n"
f"Attempted: {response.tool_call}\n"
f"Error: {error}"
)},
])
return out.text.strip()
The 25-word cap is doing real work. It forces the correction down to signal and keeps you from smuggling the contaminated trace back in under the banner of "context." A good distilled note reads like "the date argument must be ISO 8601, not a natural-language string," not like a replay of the broken call.
A retry policy that knows when to stop
Clean retries recover more often, but recovery is not guaranteed, and the worst thing you can do is loop forever on a task the agent cannot do. A retry policy ties it together: bound the attempts, distill between them, and escalate instead of spinning. Escalation might be a stronger model, a human in the loop, or a clean failure the caller can handle, but it should never be a fourth identical attempt.
def run_with_retries(task: str, max_attempts: int = 3) -> StepResult:
feedback = None
for attempt in range(1, max_attempts + 1):
result = run_step(task, feedback) # clean context each time
if result.ok:
return result
log.warning("step failed", attempt=attempt, error=result.error)
if attempt < max_attempts:
# Distill now so the NEXT clean context carries only the lesson.
feedback = distill_feedback(task, result.response, result.error)
# Out of attempts. Escalate with the distilled reason, not the raw trace.
return escalate(task, reason=feedback or result.error)
Two details matter here. The distilled feedback accumulates knowledge across attempts without accumulating tokens, because each new context carries only the latest one-line correction, not a growing pile of them. And the escalation path hands off the distilled reason, so whatever picks up the task next, a bigger model or a person, starts from the lesson rather than from the mess. If you want to go further, you can detect when two consecutive distilled notes say the same thing and escalate early, since a repeated correction means more attempts will not help.
The parts that bite
A few things are worth knowing before you ship this.
Not every failure deserves a clean reset. A transient error, a rate limit, a timeout, a 503, should be retried as-is with backoff, because the context was never the problem. Reserve clean-context retries for failures where the model's own output was wrong. Mixing the two up means you either throw away good context on a network blip or carry contamination through a real reasoning error.
Side effects do not reset when the context does. If the failed attempt already wrote to a database or sent a request, a fresh retry can repeat that side effect. Clean-context retry is about the reasoning window, not about undoing actions, so it belongs next to idempotency keys, not instead of them. See idempotency keys for retry-safe side effects for that half of the problem.
Distillation can lie. A small model summarizing a failure can drop the one detail that mattered or invent a correction that sounds right and is not. Keep the raw trace in your logs so you can audit what the distiller produced against what actually happened, the same way you would keep the full record for deterministic replay when you are debugging a nondeterministic run.
And watch the multi-step case. In a longer trajectory the contamination is not just the single failed step, it is every stale observation that step dragged in. Clean-context retry handles the immediate failure, but keeping the broader window healthy across a long run is a related discipline worth building alongside it.
The takeaway
Retrying an agent step in the context that just failed feels like giving the model the information it needs to recover. It is actually handing it the mistake and asking it not to look. The model cannot not look. Fork a clean context for every retry, carry a short distilled correction instead of the broken transcript, bound the attempts, and escalate when the correction stops changing. You recover more often and you pay less to do it, which is a rare combination in this work.
If your agent retries a lot and recovers rarely, that gap is usually contamination rather than a weak model, and it is worth fixing before you reach for a bigger one. Book a consultation call and we can look at your retry paths, measure what your contaminated attempts are actually costing you, and put clean recovery in place.
Viral Ruparel
Generative AI consultant helping teams ship reliable LLM and agent systems in production.
Contact Viral about your AI project →Frequently Asked Questions
What is context contamination in an agent retry?+
Context contamination is what happens when a failed attempt stays in the model's context window during a retry. The model sees its own earlier reasoning, the broken tool call, and the error, and it conditions the next attempt on all of it. Instead of a clean second shot the retry is a biased one, primed toward the approach that just failed, so it tends to repeat the mistake or produce a close variant of it. The failure is not in the model's ability to recover, it is in the context you handed it to recover from.
Why does retrying in the same context make things worse?+
A language model predicts the next token from everything in the window, and it cannot choose to ignore the failed attempt sitting above the retry instruction. The flawed reasoning and the failed output are now part of the prior it samples from, so the most probable continuation is something close to what it already tried. You also keep paying for those contaminated tokens on every attempt, and a long poisoned trace can push the model past the point where it reasons cleanly at all. More attempts in the same context usually means more money for a lower chance of recovery.
How do I fix a contaminated retry?+
Retry on a fresh context instead of the one that failed. Fork a clean window that contains the original task and a short distilled note of what went wrong, and leave the raw failed transcript out of it. The distilled note should say what the error was and what to avoid, not show the entire broken attempt, so the model gets the correction without the priming. Keep the failed traces in your logs for debugging, just not in the window the next attempt reasons from.
Should I ever keep the failed attempt in context?+
Sometimes a compact, accurate error message helps the model correct a narrow mistake, for example a malformed argument it can fix in one step. The problem is the full contaminated trace, not a single clear piece of feedback. The reliable pattern is to distill the failure into the smallest correct instruction that would change the next attempt, add that to a clean context, and discard the rest. If even the distilled feedback does not help after a couple of tries, that is your signal to escalate rather than keep burning attempts.
Related Articles
Your Evals Are Green and the Product Is Getting Worse
A frozen eval suite stops describing your product within weeks. New intents, new tools, new retrieval indexes and new traffic shapes all land after the set was written, so the build stays green while real users hit failures the set has never seen. This is how to measure eval-set drift against live traffic, mine novel production failures, promote them into a living golden set, and gate releases on coverage and freshness instead of a pass rate alone.
Your Agent's One Reliability Number Is Hiding the Failures That Matter
A single pass rate on a dashboard tells you an agent is "99% reliable" and tells you nothing about which 1% is on fire. Meanwhile the slow answers, the blown budgets, and the confidently wrong responses all average into the same green number. This is how to define real SLOs for an agent across the dimensions that actually break, track an error budget against each one, alert on burn rate before users feel it, and gate releases on the budget instead of on a vibe.
When Every AI Provider Goes Down at the Same Time
Early this month three of the big model providers degraded within a few hours of each other, and a lot of AI features went dark with them. If your product calls one provider and waits, their outage is now your outage. This is how to build a resilience layer in front of your model calls: a circuit breaker that fails fast, health-aware failover across providers, request hedging for the slow tail, and a degraded mode that keeps the lights on when nothing is healthy.