•11 min
Your Agent Got the Right Answer the Wrong Way
Your eval checks the final answer and goes green. Meanwhile the agent called six tools to do the work of two, hit a write endpoint it never needed, and landed on the right output by luck. Output-only evals cannot see any of that, and the newer reasoning models take longer, more autonomous paths where it matters more. Trajectory evaluation grades the steps: which tools ran, in what order, with what arguments. Here is how to build it and gate it in CI.
AI Agents
Evals