Your Agent Failed at Step 7. The Bug Was at Step 3.
When a multi-step agent produces a confidently wrong answer, the step that threw the error is almost never the step that caused it. The real culprit is an earlier decision whose bad output only surfaced downstream, so you patch the symptom and the failure comes back next week. This is agent failure attribution, and here is how to find the decisive step with counterfactual replay instead of guessing.
A team called me in because their research agent had a failure they could not kill. It ran a chain of about a dozen steps, searched, read, reasoned, called a few tools, and produced a report. Every week or two it would produce a report that was confidently, specifically wrong, and every week or two an engineer would open the trace, find the step that threw an error or looked off, patch it, and ship. The next failure always looked new. They were playing whack-a-mole with a trajectory that had a dozen places to hide.
When I looked at one of these traces the error was at step nine, a tool call that returned a result the model clearly misread. So the engineer before me had hardened step nine. But step nine was fine. The tool had returned exactly what it was asked for. The problem was step three, a retrieval step that had quietly pulled the wrong source document six steps earlier. Everything after step three was the model doing competent work on a bad premise. Step nine was just where the bad premise finally produced something that looked like an error.
This is the thing almost nobody gets right about debugging agents. The step that fails is not the step that broke your run. And until you can tell those two apart on demand, you are going to keep fixing the loud step and shipping the bug.
Why the loud step is almost never the guilty step
A single-call model either gives you a good answer or a bad one, and the bug is right there in front of you. A multi-step agent is different. It makes a sequence of decisions where each one conditions the next, so a mistake made early does not announce itself. It propagates. The model takes the bad output as a given and reasons perfectly well on top of it, which means the trajectory stays smooth and plausible right up until the moment the bad premise collides with something that checks it.
That collision is what you see in the trace. A tool returns an error, a validator rejects the output, the final answer contradicts a known fact. It feels like the root cause because it is the first thing that looks wrong. It is not. It is the first thing that looks wrong to you. The actual decision that doomed the run happened earlier, in a step that returned confidently and moved on.
The research literature has started putting a precise name on what you are actually hunting for. The term is the decisive error, defined as the earliest step whose correction would change the outcome from failure to success under a counterfactual intervention on the trajectory. That definition is worth reading twice, because it is both exactly right and the thing your instincts fight. You are not looking for the step that looks worst. You are looking for the earliest step you would have to change to make the whole run succeed. Those are different questions and they usually have different answers.
The business cost of answering the wrong question is not subtle. Every misattributed failure means an engineer spends an afternoon hardening a step that was never broken, ships a patch that narrows one symptom, and leaves the real cause live to generate the next variant. Your mean time to repair stays high, your failure rate barely moves, and your team slowly loses trust in the agent because the same class of bug keeps coming back wearing a different hat.
The approach: replay the run, change one step, see if it survives
If the decisive step is the earliest step whose correction flips the outcome, then you can find it by doing exactly that. Take the failed trajectory, pick a candidate step, replace its output with a corrected one, replay everything after it, and check whether the run now succeeds. Walk candidates from earliest to latest and the first one that flips the outcome is your decisive step. This is counterfactual replay, and it is the most reliable attribution method I know of because it tests causation directly instead of inferring it from how bad a step looks.
The one hard prerequisite is that you can replay a run deterministically from any point, feeding recorded inputs and a chosen override for the step under test. If you do not have that yet, build it first. I wrote a whole piece on deterministic replay for nondeterministic agents, and everything below assumes you can rerun a trajectory from step k with the earlier steps pinned to what actually happened.
Here is the shape of the trap first, so the fix lands. This is the attribution logic most teams actually run, even if they would not write it down this way.
# The naive attribution: blame the step where the error showed up.
# This is what "open the trace and find the first thing that looks wrong" does.
def naive_attribution(trace):
for step in trace.steps:
if step.raised_error or step.output_flagged_by_validator:
return step # the LOUD step, not the guilty one
return trace.steps[-1] # nothing obvious, blame the final answer
The problem is baked into the loop. It returns the first step that looks wrong, which is by construction the step where a buried problem finally surfaced. It has no way to see that the step three outputs upstream was the one that mattered, because that step did not look wrong at all.
Now the counterfactual version. The core idea is a single function that asks, for one candidate step, whether fixing it would have saved the run.
# Replay the trajectory with one step's output replaced, keeping everything
# before it exactly as it originally ran. Returns True if the run now succeeds.
def correction_flips_outcome(trace, step_index, corrected_output, verify):
# Pin steps [0, step_index) to their recorded outputs, so we isolate the
# effect of this one change and nothing upstream drifts.
overrides = {i: trace.steps[i].output for i in range(step_index)}
overrides[step_index] = corrected_output
# Replay from the start using the overrides; steps after step_index run live
# and get to react to the corrected output.
result = replay(trace, overrides=overrides)
# verify() is your success check for this task: an assertion, a golden
# answer match, or an LLM judge. It defines what "the run succeeded" means.
return verify(result)
With that primitive, attribution is a walk from the earliest step to the latest, stopping at the first correction that flips the outcome.
def find_decisive_step(trace, propose_correction, verify):
"""Return the earliest step whose correction turns failure into success."""
for i, step in enumerate(trace.steps):
# Ask a strong model (or a human, or a known-good oracle) what this step
# SHOULD have produced, given the context it actually saw.
corrected = propose_correction(step, context=trace.steps[:i])
if corrected is None:
continue # no plausible better output; this step was likely fine
if correction_flips_outcome(trace, i, corrected, verify):
return i # earliest flip = decisive step
return None # no single-step correction recovers the run; see pitfalls
That is the whole method. You are no longer asking which step looks worst. You are asking, step by step from the beginning, which is the first one where a better output would have saved the run. The first yes is the decisive step, and it is the thing you actually fix.
Making it cheap enough to run in anger
The obvious objection is cost. A naive scan replays the trajectory once per candidate step, and each replay burns real tokens and real tool calls. On a twelve step trace that is twelve replays, and on anything longer it gets expensive fast. Two things make this practical.
First, you do not need to test every step. The decisive step has a monotonic property that is close enough to true to exploit: if correcting step k does not recover the run, correcting any step before k usually will not either, because the damage compounds forward. That lets you binary search the trajectory instead of scanning it, which turns a dozen replays into three or four.
def find_decisive_step_bsearch(trace, propose_correction, verify):
lo, hi = 0, len(trace.steps) - 1
decisive = None
while lo <= hi:
mid = (lo + hi) // 2
corrected = propose_correction(trace.steps[mid], context=trace.steps[:mid])
if corrected and correction_flips_outcome(trace, mid, corrected, verify):
decisive = mid # this one flips; look earlier for the ROOT
hi = mid - 1
else:
lo = mid + 1 # no flip here; the cause is later
return decisive
Second, use a cheap critic to rank candidates before you spend a replay on them. Before running any counterfactual, have a small model score each step for how likely it is to be the first real mistake, given only the context that step had available. Test the high-scoring candidates first. The critic is often wrong about the exact step, which is why you still verify with replay, but it is good enough to reorder the search so your first real replay lands near the answer. This is the same division of labor I lean on everywhere: a cheap model proposes, an expensive check disposes. It pairs naturally with grading the whole trajectory rather than only the final answer, since the trajectory grader already produces per-step signal you can feed the critic.
The parts that bite
Counterfactual attribution is powerful, and it has sharp edges worth knowing before you wire it into your on-call flow.
The correction has to be honest. The whole method rests on propose_correction returning what the step genuinely should have produced given the context it actually had. If your corrector cheats by using information the step could not have known, you will attribute the failure to a step that was doing the best possible job with what it was given, and you will miss the real upstream cause. Constrain the corrector to the context prefix, the same way the original step was constrained.
Side effects do not replay cleanly. If a step in the middle of the trajectory wrote to a database or sent a message, replaying from before it can repeat that side effect or, worse, hit a world that has since changed. Attribution belongs on recorded traces in a sandbox where tool calls are served from the recording, not re-executed against production. Keep the raw trace, replay against it, and never let a counterfactual run touch a live system.
Some failures have no single decisive step. Occasionally two separate mistakes each half-break the run, and no single correction recovers it because fixing one leaves the other. When find_decisive_step returns nothing, that is the signal, not a bug. It means you are looking at a multi-cause failure and you need to search for the smallest set of steps whose joint correction recovers the run, which is a harder problem but a rarer one. Do not let the common single-cause case wait on the rare multi-cause case.
And attribution tells you where, not why. Finding that step three pulled the wrong document is the start of the fix, not the end. You still have to figure out whether the retrieval query was bad, the index was stale, or the ranking was off. Attribution's job is to point you at the one step that is worth that investigation, so you stop spending the investigation on step nine.
The takeaway
A wrong answer from a long agent run is a crime scene, and the step that threw the error is just the body. The decision that actually killed the run happened earlier, in a step that returned confidently and looked completely fine. If you debug by opening the trace and fixing the first thing that looks wrong, you will harden the symptom and ship the cause, over and over, which is exactly the loop that makes a flaky agent feel unfixable.
The way out is to stop trusting how steps look and start testing what they caused. Record your runs so you can replay them, then for a failed run find the earliest step whose correction flips the outcome. Binary search it, rank candidates with a cheap critic, and verify every candidate with a real replay. You will spend a few extra calls per investigation and get back the one thing that actually matters, which is the single step you need to fix.
If your agent keeps failing in ways that feel new every time but smell the same, you are almost certainly fixing loud steps instead of guilty ones. Book a consultation call and we can set up replay and counterfactual attribution on your trajectories, find the steps that are really costing you, and get your repairs landing on the cause for once.
Viral Ruparel
Generative AI consultant helping teams ship reliable LLM and agent systems in production.
Contact Viral about your AI project →Frequently Asked Questions
What is failure attribution in an LLM agent?+
Failure attribution is the process of taking a failed agent run and identifying the specific step that actually caused the failure, not just the step where the error became visible. In a multi-step or multi-agent system the wrong final answer is usually the result of a bad decision made several steps earlier, which stayed silent until a later step acted on it. Attribution answers the question of which step you would have to change to make the run succeed, so you fix the cause instead of the symptom.
What is the decisive step or decisive error?+
The decisive step is the earliest step in the trajectory whose correction would flip the outcome from failure to success. Formally it is the earliest agent and step pair where, if you replaced its output with a correct one and let the rest of the run proceed, the final result would have been right. It is the root cause in a form you can act on, because it points at the one decision that mattered rather than at every step that looked a little off.
How does counterfactual replay find the decisive step?+
Counterfactual replay reruns the trajectory from a candidate step with that step's output replaced by a corrected version, while keeping everything before it fixed. If the corrected run succeeds, that step was decisive or close to it. By testing candidate steps from earliest to latest, or with a binary search over the trajectory, you find the earliest step whose correction changes the outcome. It depends on being able to replay a run deterministically from any point.
Why not just blame the step that threw the error?+
Because the step that errors is where the failure surfaced, not where it started. A retrieval step that pulled the wrong document does not error, it returns confidently, and the failure only appears three steps later when a tool call or a final answer is built on that bad context. If you patch the visible error you harden the symptom while the real cause keeps producing new variants of the same failure. Attribution exists precisely because the loud step and the guilty step are usually different.
Related Articles
Your Agent Retries Are Making the Failure Worse
When an agent step fails and you retry it in the same context, the failed attempt stays in the window and the model conditions its next try on its own mistake. The retry does not get a fresh shot, it gets a biased one, and it often repeats or compounds the error while you pay for every token. This is how context contamination kills retry recovery, and how to fix it with clean context forks, distilled feedback, and a retry policy that knows when to stop.
Your Evals Are Green and the Product Is Getting Worse
A frozen eval suite stops describing your product within weeks. New intents, new tools, new retrieval indexes and new traffic shapes all land after the set was written, so the build stays green while real users hit failures the set has never seen. This is how to measure eval-set drift against live traffic, mine novel production failures, promote them into a living golden set, and gate releases on coverage and freshness instead of a pass rate alone.
Your Agent's One Reliability Number Is Hiding the Failures That Matter
A single pass rate on a dashboard tells you an agent is "99% reliable" and tells you nothing about which 1% is on fire. Meanwhile the slow answers, the blown budgets, and the confidently wrong responses all average into the same green number. This is how to define real SLOs for an agent across the dimensions that actually break, track an error budget against each one, alert on burn rate before users feel it, and gate releases on the budget instead of on a vibe.