Your Agent Got the Right Answer the Wrong Way
Your eval checks the final answer and goes green. Meanwhile the agent called six tools to do the work of two, hit a write endpoint it never needed, and landed on the right output by luck. Output-only evals cannot see any of that, and the newer reasoning models take longer, more autonomous paths where it matters more. Trajectory evaluation grades the steps: which tools ran, in what order, with what arguments. Here is how to build it and gate it in CI.
A team I worked with last month had a support agent that passed its whole eval suite the morning they asked me to look at it. Green across the board. The same agent, in production, was quietly calling their refund tool on read-only questions. It never issued a bad refund, because a downstream permission check caught it every time, so the final answers stayed correct and the eval stayed green. The agent was reaching the right destination through a door it should never have touched, and nothing in the test harness could see the door.
That is the blind spot in most agent evals. We grade the answer. We almost never grade the path. And with the reasoning models that shipped this month taking longer, more autonomous multi-step routes before they respond, the path is where the cost, the latency, and the sharp edges now live.
What an output-only eval cannot see
Think about what a typical agent test asserts. You feed a task, let the agent run, and check the final output against an expected value or a rubric. If the output matches, pass. That tells you the agent can be right. It tells you nothing about how.
Here is the same task, two runs, both passing an output check:
- Run A: calls
search_docs, reads the top result, answers. Two steps. - Run B: calls
search_docs, thensearch_docsagain with the same query, thenget_account, thenlist_invoices, thenissue_refund(rejected by permissions), then finally answers from the first search result. Six steps, one of them a write it had no business making.
Run B is a failure wearing a passing grade. It costs three times the tokens, it is slower, it pinged a dangerous tool, and it got the right answer by luck rather than by process. The day the permission check has a gap, or the day a slightly different input makes the refund look valid, Run B is an incident. Your eval will still be green that morning.
This is why leaderboard scores and internal pass rates keep disagreeing with what teams actually see in production. The score measures whether the answer was reachable. Production measures the route the agent actually took, every time, under inputs you did not test. Trajectory evaluation closes that gap by making the route a first-class thing you grade.
The trajectory is already in your traces
You do not need new instrumentation to start. If you have any observability on your agent, you are already recording the raw material: the ordered tool calls, their arguments, and what came back. The first step is to pull that into a clean structure you can assert against.
from dataclasses import dataclass
@dataclass
class Step:
tool: str
args: dict
ok: bool # did the call succeed, from the tool's own result
def extract_trajectory(messages: list[dict]) -> list[Step]:
"""Flatten an agent run into an ordered list of tool calls.
Works off the standard tool-call message shape most SDKs emit."""
steps: list[Step] = []
for msg in messages:
for call in msg.get("tool_calls", []):
steps.append(Step(
tool=call["function"]["name"],
args=call["function"]["arguments"], # already a dict here
ok=True, # patched below when we see the matching result
))
return steps
Once a run is a list of Steps, everything downstream is ordinary code. That is the whole trick: turn the agent's behavior into data, then assert on the data the way you would assert on anything else.
Start with invariants, not a golden path
The instinct is to record one perfect run and diff every future run against it. Resist that first. A single golden trajectory is brittle. The moment the agent finds a legitimately better two-step path, your eval fails a genuine improvement, and you learn to ignore red, which is worse than having no eval at all.
Begin with invariants instead. These are properties a good trajectory must have, regardless of the exact steps, and they catch the expensive failures without pinning the agent to one route.
def check_invariants(steps: list[Step], spec: dict) -> list[str]:
"""Return a list of violations. Empty means the path is acceptable.
spec declares what a good trajectory must and must not do for this task."""
violations = []
called = [s.tool for s in steps]
# required tools must appear at least once
for tool in spec.get("must_call", []):
if tool not in called:
violations.append(f"missing required tool: {tool}")
# forbidden tools must never appear (e.g. no writes on a read task)
for tool in spec.get("must_not_call", []):
if tool in called:
violations.append(f"called forbidden tool: {tool}")
# step budget: catches loops and wasteful detours
budget = spec.get("max_steps")
if budget is not None and len(steps) > budget:
violations.append(f"used {len(steps)} steps, budget was {budget}")
# no immediate repeat of the same call with the same args
for a, b in zip(steps, steps[1:]):
if a.tool == b.tool and a.args == b.args:
violations.append(f"repeated identical call: {a.tool}")
return violations
For the refund case at the top of this piece, the spec is one line of intent: on a read-only question, issue_refund goes in must_not_call. That single invariant would have turned the green eval red on the day the behavior started, months before a customer ever felt it. Invariants are cheap, they are deterministic, and they never flake, which makes them the right thing to gate a build on.
If you are already collecting spans through OpenTelemetry or a similar setup, the extraction step is mostly reading fields you have, and the invariant checks slot straight onto the trace. I go deeper on getting that signal out of a running agent in the piece on agent observability with OpenTelemetry.
When several good paths exist, grade with a rubric
Invariants handle the hard constraints. They do not capture softer quality, like whether the agent chose a sensible tool over a merely allowed one, or whether the order was logical. For open-ended tasks where you cannot enumerate every acceptable route, add a rubric graded by a judge model. The key is to make the judge score the process, and to feed it the trajectory rather than just the answer.
import json
from openai import OpenAI
client = OpenAI()
JUDGE = """You grade the PROCESS an agent used, not its final answer.
You are given the task, the ordered tool calls with arguments, and a rubric.
Score each rubric item 0 to 1 and return JSON:
{"scores": {"<item>": 0.0}, "notes": "one line per weak item"}
Judge only what the steps show. Do not reward a good answer reached by a bad path."""
def judge_trajectory(task: str, steps: list, rubric: list[str]) -> dict:
path = "\n".join(
f"{i+1}. {s.tool}({json.dumps(s.args)})" for i, s in enumerate(steps)
)
resp = client.chat.completions.create(
model="claude-sonnet-5-5",
messages=[
{"role": "system", "content": JUDGE},
{"role": "user", "content":
f"Task:\n{task}\n\nTrajectory:\n{path}\n\nRubric:\n" +
"\n".join(f"- {r}" for r in rubric)},
],
response_format={"type": "json_object"},
temperature=0,
)
return json.loads(resp.choices[0].message.content)
# rubric = [
# "used a search tool before answering rather than guessing",
# "did not fetch data unrelated to the question",
# "reached the answer in the fewest reasonable steps",
# ]
A judge that sees the path can tell you the agent retrieved before answering, or that it wandered into three unrelated lookups first. A judge that sees only the answer cannot. Judges have their own failure modes, and grading a trajectory inherits all of them, so treat the score as a calibrated signal rather than truth. If you have not tightened up your judges yet, the pitfalls I cover in LLM-as-judge bias and calibration apply directly here.
Wire it into CI as a gate
The point of all this is a build gate that fails when the path degrades, not a dashboard someone glances at once a quarter. Combine the two signals: deterministic invariants block the build outright, and the aggregate judge score has to clear a threshold with a margin.
def evaluate_case(agent_run, spec: dict, rubric: list[str]) -> dict:
steps = extract_trajectory(agent_run["messages"])
violations = check_invariants(steps, spec)
judged = judge_trajectory(agent_run["task"], steps, rubric)
quality = sum(judged["scores"].values()) / max(len(judged["scores"]), 1)
return {
"passed": not violations and quality >= 0.8,
"violations": violations, # any entry is a hard fail
"quality": round(quality, 2), # soft signal, thresholded
"steps": len(steps), # track this trend over time
"notes": judged.get("notes", ""),
}
def run_suite(cases) -> bool:
results = [evaluate_case(**c) for c in cases]
failed = [r for r in results if not r["passed"]]
avg_steps = sum(r["steps"] for r in results) / len(results)
print(f"{len(results) - len(failed)}/{len(results)} passed, "
f"avg {avg_steps:.1f} steps/case")
for r in failed:
print(f" FAIL q={r['quality']} steps={r['steps']} {r['violations']} {r['notes']}")
return not failed # exit non-zero in CI when this is False
Track average steps per case as a first-class metric across builds, not just pass or fail. A suite that stays green while the step count creeps from four to nine is telling you the agent is getting slower and more expensive one prompt tweak at a time. That trend line is often the earliest warning that a change quietly made the agent worse, and it is invisible to any output-only check.
This slots in beside your existing answer-quality suite rather than replacing it. If you already run multi-turn conversation evals, the trajectory checks layer straight onto the same recorded runs, and the two together give you a far honest picture than either alone. I wrote about the conversation side of that in multi-turn agent evaluation with a user simulator.
Tradeoffs and the parts that bite
Golden trajectories rot fast. If you do pin exact paths for your most deterministic cases, expect to update them whenever the agent legitimately improves, and keep that set small. Lean on invariants and rubrics for everything else so a better path reads as a pass, not a regression.
Judges cost money and add variance. Grading every step of every case with a model gets expensive, and a single judge call is noisy. Run the deterministic invariants on every case, since they are nearly free, and sample the judge on the cases that need the softer signal. Run each judged case a few times and average, so one bad draw does not fail a build.
Invariant specs are real work to write. Every task needs someone to say what a good path looks like, which tools are required, which are off limits, what the step budget is. That is a feature, not a tax. The act of writing it forces a conversation about what the agent is actually allowed to do, and that conversation surfaces the refund-tool problem long before the code does.
Do not gate on the judge alone. An LLM judge grading a trajectory can be talked into a high score by a plausible-looking path, the same way it can be swayed by a long answer. Deterministic invariants are the floor that a clever wrong path cannot argue its way past. Keep them as the hard gate and let the judge shape the soft score above it.
The takeaway
An agent that reaches the right answer through the wrong path is not passing, it is failing quietly with a green light on. Output-only evals were built for a world where the model answered in one shot, and that world is gone. Grade the trajectory: extract the ordered tool calls, enforce the hard invariants deterministically, score the softer quality with a judge that sees the path, and gate CI on both while you watch the step count trend. None of it requires new infrastructure, only the decision to treat how the agent got there as something worth checking.
If your agents pass every eval and you still get surprised in production, the gap is almost always in the path, and it is measurable in an afternoon. Book a consultation call and we can look at where your agents are getting the right answers the wrong way, and put a gate in front of it.
Viral Ruparel
Generative AI consultant helping teams ship reliable LLM and agent systems in production.
Contact Viral about your AI project →Frequently Asked Questions
What is trajectory evaluation for an AI agent?+
Trajectory evaluation grades the sequence of steps an agent took to reach its answer, not only the answer itself. A trajectory is the ordered list of tool calls, their arguments, and the observations that came back. You score that path against either a reference trajectory or a rubric, checking whether the agent used the right tools, in a sensible order, with correct arguments, and without unnecessary or unsafe calls. It catches the case where an agent produces a correct output through a wasteful or dangerous route that an output-only eval reports as a pass.
Why is checking the final answer not enough for agents?+
Because an agent can be right for the wrong reasons. It can call a write endpoint it never needed, loop on a tool five times, retrieve from the wrong source and get lucky, or take a ten-step path where two steps would do. All of those pass an output-only assertion and all of them are latent failures: they cost more, they are slower, and they break the moment the input shifts slightly. Grading the path surfaces these before a customer does.
Do you need a golden trajectory for every test case?+
No. A golden reference trajectory is the strictest option and it works well for narrow, deterministic tasks. For open-ended tasks where several good paths exist, use a rubric graded by an LLM judge, or check invariants instead: required tools were called, forbidden tools were not, no unnecessary writes happened, and the step count stayed under a budget. Invariant checks are cheap, deterministic, and catch most of the damage without pinning the agent to one exact route.
How do you run trajectory evals in CI without them being flaky?+
Score structural invariants deterministically so they never flake, and reserve the LLM judge for the softer quality signal with a margin built into the threshold. Pin model versions, seed where you can, and run each case a few times to measure variance rather than trusting a single run. Gate the build on the deterministic checks and the aggregate judge score, and treat a single noisy judge call as data, not a verdict.
Related Articles
Your Agent Would Rather Guess Than Admit It Doesn't Know
Most agents never say "I don't know." They were trained on benchmarks that reward a confident guess over an honest abstention, so in production they invent a policy, a number, or a citation and say it with a straight face. An abstention gate fixes that: a cheap confidence signal, a threshold you calibrate against a real error budget, and a decision to answer, defer, or escalate. This is how you get an agent that knows its own limits.
Nobody Can Tell You How Many Agents You Are Running
A year of shipping agents leaves most teams with a fleet nobody can inventory: no owner, no declared budget, no record of which tools each one can reach. That is agent sprawl, and it shows up as a surprise invoice and an incident with no name on it. The fix is boring and it works: an agent registry. A typed manifest every agent must declare, a CI gate that blocks anything unregistered or over-budget, and a reconciliation job that flags an agent doing more than it said it would.
Your Agent's Memory Is Full and Most of It Is Junk
You gave your agent persistent memory and it worked. Six months later the store is full of duplicates, contradictions, and vague paraphrases that crowd out the facts you actually need, and retrieval quietly gets worse every week. This is a garbage-collection problem, not a storage problem. Here is how to build the curation layer: gate what gets written, deduplicate and reconcile on the way in, decay what stops earning its slot, and measure whether the store is still healthy.