LLM Evals in Production: From a CI Gate to a Living Eval Set

How to build LLM evals that hold up in production: a CI gate, a calibrated judge, agent tests that see the path and the conversation, a set that feeds itself from traffic.

Guide10 essays21 min readUpdated

Most teams I work with arrive at evals the same way. Something broke in production that no test could have caught: a prompt tweak that made the model cautious across the board, a model bump that changed how tool calls were formatted, an embedding swap that quietly returned the wrong chunks. Someone says "we should have evals," a few dozen cases get written in an afternoon, a pass rate shows up in CI, and everyone relaxes. Then the number stays green while the product gets worse, and the team learns the second lesson: a pass rate is only as good as the cases behind it, the judge scoring them, and the discipline that keeps both honest.

This guide is the calm version of ten essays I have written about that second lesson. Each takes a single failure, the judge that prefers long answers, the agent that scores 90% and still fails customers, the set that aged off the product, and works through the fix in code. Together they describe one system: a gate in CI, graders checked against humans, agent tests that see the conversation and the path and not only the final text, a set that feeds itself from production, and the same gate reused for every change that can move quality, whether a prompt, a model id or an index.

Nothing here needs new infrastructure. It is a golden dataset in git, a scorer, a script that exits nonzero, and the habit of running it before anything reaches a user.

Start with a gate in CI, not a dashboard

Observability tells you what already happened to real users. It is a rear-view mirror, and by the time a refusal rate creeps up on a chart the regression shipped days ago. An eval runs a fixed set of inputs through your prompt or agent before anything reaches a user and scores the outputs against a bar you set. It is a unit test for something fuzzy, so you cannot assert that the output equals an expected string. You score how good it is.

Three parts. First, a golden dataset: twenty to fifty cases that represent your real traffic and, more importantly, your real failure modes, kept in version control so it reviews like code. Second, a scorer. Every case carries cheap deterministic checks, must_include and must_not_include, plus a plain-language rubric for what a substring match cannot capture: tone, correctness, whether the request was resolved. Run the assertions first, because they are free and unambiguous, and a hard assertion failure is a zero however nice the answer reads. Only then does a separate model grade the rubric. Third, the gate: aggregate, compare against a threshold, exit nonzero so the pull request cannot merge.

def score_case(system_under_test, case: dict) -> dict:
    output = system_under_test(case["input"])
    hard_failures = check_assertions(output, case)   # must_include / must_not_include
    if hard_failures:
        return {"id": case["id"], "score": 0, "reason": "; ".join(hard_failures)}
    verdict = judge(output, case)                      # separate model, strict rubric
    return {"id": case["id"], "score": verdict["score"], "reason": verdict["reason"]}

def main() -> int:
    results = [score_case(answer, case) for case in CASES]
    avg = mean(r["score"] for r in results)
    zeros = [r for r in results if r["score"] == 0]
    if avg < MIN_MEAN_SCORE or len(zeros) > MAX_ZEROS:
        print(f"FAILED: {len(zeros)} hard failures, mean {avg:.2f}")
        return 1
    print("PASSED")
    return 0

Gate on the floor as well as the average. A mean of 4.2 looks healthy while your three most important cases sit at 2, so add per-case minimums for anything business-critical. Leave a buffer between the threshold and where cases actually score, because models vary run to run and a case that wobbles across the line on identical code sends you chasing a ghost. Keep the pull request gate to a tight, fast core and push the exhaustive sweep to a nightly job, because a gate people wait for beats one they route around. I walked through the whole build, GitHub Actions included, in LLM evals in CI.

Make the judge earn its number

Once a model grades your outputs, the judge becomes the load-bearing measurement in the system. It gates releases, picks winners in prompt comparisons, and increasingly decides which outputs become training data. If its number is biased, the bias propagates into every decision downstream, quietly, because a biased judge does not throw. It returns a plausible score with a plausible reason that reads ten percent high.

Three biases do most of the damage. Position bias: in a pairwise comparison the judge prefers whichever answer it saw first, by a double-digit margin on identical answers. Run every comparison in both orders and only count a verdict that survives the swap. If the judge flips when you flip the order, record a tie, because it was voting on position. Verbosity bias: a correct three-sentence answer loses to the same answer padded to three paragraphs, often by fifteen to thirty points. Tell the judge not to reward length, then make the bias visible by tracking the correlation between score and word count across the set. Self-preference bias: a judge rates outputs in its own style higher, by ten to twenty-five percent, so judge with a different model family than the system under test.

None of that tells you whether the judge's numbers track reality. For that you need calibration: a few hundred examples graded by hand, spanning easy cases and known failure modes. Run the judge over the same set, bucket both into pass or fail at your gate threshold, and measure agreement with Cohen's kappa, not correlation. Above about 0.6 the judge agrees with your humans well enough to gate on. Below that, the disagreements are the queue to route to a person. Re-label a fresh sample whenever the judge model, the rubric or the traffic changes, and pin the judge version so a provider update cannot move your baseline without a deploy. The fixes and the calibration code are in calibrating an LLM judge you can trust.

Test the conversation, not a single prompt

A single-shot eval sends one input, gets one output, scores it. That is the right shape for a summarizer or a classifier. A conversational agent fails in the gaps between turns: it honors a constraint for two turns and forgets it on the third, asks in turn five for the order number the user gave in turn one, or caves under pushback and issues a refund outside the policy window. None of those exist inside a single prompt, and a scripted user cannot follow an agent whose next question depends on what it said last.

The fix, borrowed from tau-bench and its successors, is to put another model on the other side of the conversation. A scenario defines the persona, the goal, the facts the user reveals only when asked, the starting state of a sandboxed backend, and two rubrics: what success looks like and what the agent is not allowed to do. The agent under test is the real thing, with its tools bound to the sandbox. A loop alternates simulated user and agent until the simulator signals it is done or a turn cap fires.

Two choices decide whether the suite finds bugs or rubber-stamps everything. The first is an adversarial simulator. A helpful model volunteers the order number, accepts the wrong cancellation, and ends the conversation happy. Write personas that withhold, push back, and occasionally act in bad faith. If every scenario passes on the first run, the simulator is too nice. The second is scoring on two axes, separately. Task success comes from the final state of the sandbox, deterministic and free, because the transcript is only what the agent claims and an agent will narrate a refund it never processed. Policy adherence needs a judge reading the whole transcript. An agent that completes the task by breaking a rule is a partial fail you need to see.

Pin model versions, keep the simulator temperature low, cap the turns, and run the money paths and safety policies on every pull request with the full persona sweep nightly. The scenario, sandbox, simulator, loop and scorer are all in multi-turn agent evaluation with a user simulator.

Measure consistency, not an average

A 90% pass@1 score hides two lies at once. The first is variance. pass@1 runs each case once, so it cannot distinguish an agent that passes 180 cases every time and fails 20 every time from an agent that passes every case 90% of the time. The first is a product. The second is a slot machine, and every customer pulls the lever once. The metric that captures this is pass^k, the probability that the agent succeeds on all k attempts. It is the pessimist's version of pass@k, which was designed for code generation, where you keep whichever candidate compiles. Your customer gets one run.

Run each case k times at production temperature and report two numbers: the mean per-run success rate, which is the comforting pass@1, and the fraction of cases that passed on all k runs. The gap between them is the size of the first lie, and the cases in the gap are named for you. I have watched a suite report 0.91 on the first and 0.68 on the second, and those 23 points were the support queue.

The second lie is compounding. An agent is a pipeline: understand, pick a tool, fill arguments, read the result, answer. If each step succeeds independently 90% of the time, a five-step task succeeds about 59% of the time and a ten-step task about 35%, with no bug introduced anywhere. So estimate a per-step success rate from the trajectories, because in a multiplicative system the weakest step dominates and that is where every hour should go. Argument filling at 0.83 capping a task at 0.72 is a far more useful finding than "task success is 72%."

Do not trust a small sample. Three passes out of three has a Wilson lower bound near 0.44, barely better than a coin. Report the lower confidence bound per case rather than the raw fraction, spend the sample budget on cases near the threshold, and make it the gate CI exits on:

def gate(results, min_consistency: float = 0.85, min_case_floor: float = 0.90):
    # Fraction of cases that passed on every one of their k runs.
    consistently_correct = sum(r.pass_pow_k for r in results) / len(results)
    # Any case whose honest floor is below the bar blocks the release on its own.
    weak = [
        r.case_id for r in results
        if wilson_lower_bound(r.passes, r.k) < min_case_floor
    ]
    ok = consistently_correct >= min_consistency and not weak
    if not ok:
        print(f"BLOCKED: consistency {consistently_correct:.2%} "
              f"(min {min_consistency:.0%}); weak cases: {weak}")
    return ok

The judge's own variance leaks into these numbers, which is one more reason to calibrate it first. The full argument, with the pass^k, per-step and Wilson code, is in agent reliability with pass^k and per-step consistency.

Grade the path, not only the answer

An output-only eval tells you the agent can be right. It tells you nothing about how. Two runs of the same task can both pass an answer check while one took two steps and the other took six, called the same search twice, and hit a refund endpoint on a read-only question that only a downstream permission check stopped. The second run is a failure wearing a passing grade: three times the tokens, slower, and one permission gap from an incident. With reasoning models taking longer, more autonomous routes, the path is where the cost and the sharp edges live.

The trajectory is already in your traces. Flatten the run into an ordered list of tool calls with arguments and outcomes, and everything downstream becomes ordinary code. Resist recording one perfect run and diffing against it. A golden trajectory is brittle: the moment the agent finds a legitimately shorter path, the eval fails an improvement and the team learns to ignore red. Start with invariants, properties a good path must have regardless of the exact steps.

def check_invariants(steps: list[Step], spec: dict) -> list[str]:
    """Return violations. Empty means the path is acceptable."""
    violations = []
    called = [s.tool for s in steps]
    for tool in spec.get("must_call", []):          # required tools appear
        if tool not in called:
            violations.append(f"missing required tool: {tool}")
    for tool in spec.get("must_not_call", []):      # no writes on a read task
        if tool in called:
            violations.append(f"called forbidden tool: {tool}")
    budget = spec.get("max_steps")                   # catches loops and detours
    if budget is not None and len(steps) > budget:
        violations.append(f"used {len(steps)} steps, budget was {budget}")
    for a, b in zip(steps, steps[1:]):               # no identical repeat
        if a.tool == b.tool and a.args == b.args:
            violations.append(f"repeated identical call: {a.tool}")
    return violations

For the refund case, the spec is one line of intent: issue_refund in must_not_call on a read-only question. Invariants are cheap, deterministic, and never flake, which makes them the right thing to block a build on. Where several good paths exist, add a rubric graded by a judge that sees the trajectory rather than the answer, and keep it as a soft score above the hard floor, because a judge can be talked into a high score by a plausible-looking path. Track average steps per case across builds too, because a green suite whose step count creeps from four to nine is an agent getting slower one prompt tweak at a time. The extraction, invariants, rubric and CI runner are in trajectory evaluation: grade how your agent got there.

Keep the eval set alive

A traditional test suite can sit still because the thing it tests sits still. An LLM product is a moving target of prompts, tool schemas, retrieval corpora and a model, serving traffic that moves on its own schedule. Every release ages the set: a new feature adds intents with no cases, a prompt rewrite moves the weakness your adversarial cases were testing, a reindex makes grounding cases pass trivially. All of it pushes the pass rate up while real quality goes down, so a set written in the spring scores a summer product 94% and says ship.

Measure the drift first. Embed a sample of recent production inputs and your eval inputs in the same space, cluster the production side, and check how many clusters have an eval case nearby. Weight the result by traffic: a set can cover 38 of 40 clusters and still miss a third of your volume if the two uncovered clusters are where users actually are. That number falls weeks before the support tickets rise.

Then mine production for what is missing. You rarely have ground truth on live traffic, so stack the weak signals you already emit: a thumbs-down, an online judge score below a bar, a tool error, repeated rephrasing, an escalation. Add novelty, the distance from a trace to its nearest eval case, because failures show where you are weak and novelty shows where you are blind. That cuts thousands of daily traces to the few dozen worth a human's attention. Promote representatives of clusters rather than individual traces, so each failure pattern contributes a case or two instead of forty near-duplicates of this week's incident, and land the case in the same pull request as the fix, failing before and passing after.

Finally, gate on the set's health, not only its pass rate. Traffic coverage below a floor, median case age above a limit, or too many releases since a case was added should each fail the build. The pass rate answers "did we break a known case." The health gate answers "do we still know enough cases for that to matter." Keep a human in the promotion decision and a hand-labeled slice the judge never touches, scrub PII before anything lands in the set, cap contributions per cluster, and retire cases that no longer match traffic. The coverage, mining, promotion and health-gate code is in eval-set drift and the living eval set.

Let the eval set tune the prompt

Hand-tuning a prompt is a human doing badly, from memory, the one job a search does well: try a change, measure it against everything, keep it only if it wins. An engineer can test maybe ten wordings in an afternoon, judges on the five cases on screen, and never sees the quiet regression on the other side of the distribution. The result is a prompt with contradictory rules that nobody dares refactor.

Reflective optimization replaces that loop. A reflection model reads the prompt, the cases it got wrong, and why, and proposes a targeted rewrite. The prerequisite is a metric that explains itself. A score of 0.0 teaches the optimizer almost nothing. Feedback that says the prompt confused a charge dispute with a billing question because it over-weights one word teaches it exactly what to change. DSPy's GEPA optimizer wires that metric into a search for you, and the framework-free version fits on a page: keep a small pool of candidates, expand the best, collect its failures with feedback, ask the reflection model for a child, score it, trim the pool. The output is a prompt file you can read, diff and check in, so a senior engineer can still veto a rewrite that games the metric.

The optimizer is the easy part. The gate keeps it honest. Optimize on a training split, select on validation, prove the final number on a test split you touch once, then make the comparison a build step: the candidate is promoted only if it beats the champion by a margin larger than the set's noise and does not cost more per correct answer. If your score comes from an LLM judge, the optimizer will find and exploit every bias that judge has, because exploiting the judge is what you told it to maximize. Calibrate first, read the winning prompt with your own eyes, and keep the set fresh. The loop and the CI check are in automated prompt optimization with an eval gate.

Gate every model and index change the same way

The pull request that bumps a model id is three characters long and passes review in eight seconds, and it is the most behavior-defining dependency change in your stack. Tool-call formatting shifts, verbosity and cost climb, an instruction the old model ignored is suddenly obeyed, and the prompt you tuned against the old model's quirks is now tuned against a ghost. None of it throws. Treat a model version like a database engine upgrade: pinned behind an adapter that maps a role like chat or extraction to an exact version, never a floating -latest alias; gated on an eval before promotion; rolled out through a shadow and a canary with a rollback that is a config number, not a build.

The gate runs on your own traffic. Sample a few hundred real requests with their current outputs, replay them through the candidate, and check three things at once: average quality held its bar, structured output stayed parseable on nearly every sample, and cost did not climb past a budget you set on purpose, say twenty percent, beyond which a person decides. Put tool-call format and success rate in that gate explicitly, because a model that is better at prose can be worse at picking the right tool with the right arguments. The adapter, gate and rollout code are in upgrading an LLM model without breaking production.

An embedding model swap looks even smaller and hides even better. Query vectors from the new model and document vectors from the old one live in different coordinate systems, and cosine similarity between them is a number the database will happily rank and that means nothing, so retrieval returns the wrong chunks and nothing pages anyone. The swap is a data migration. Tag every vector with the model and dimension that produced it, one namespace per version so two spaces can never blend. Backfill the new namespace idempotently while the old one keeps serving. Gate the cutover on recall@k and MRR@k over a labeled set of queries and the chunk ids that should come back, requiring the new index to match or beat the old on your corpus, whatever the public benchmark said. Flip a pointer so rollback takes seconds. Migrate one variable at a time, never chunking and embeddings together. The four rules are in migrating embedding models without breaking retrieval.

Check the answer against the evidence

Retrieval and grounding are different failure surfaces. Retrieval decides whether the right evidence made it into the context. Grounding decides whether the answer says only what that evidence supports. You can have perfect retrieval, the exact paragraph at rank one, and the model still writes thirty days where the chunk says fourteen, then staples the real citation to the claim as proof. "Only answer from the provided context" is an instruction, not a check, and "mostly follows it" is not a property you can put in front of customers.

A grounding gate verifies the output mechanically after generation. Decompose the answer into atomic claims, preserving any citation marker each sentence carried, because groundedness is a property of claims, and a paragraph-level "is this supported" check pattern-matches on topical overlap and says yes. Then check each claim against the chunk it cited, or against all retrieved chunks if it cited none, with a small fast model run in parallel. Entailment on one short claim is a narrow task cheap models do well, and the fan-out lands in the few hundred millisecond range. Hold a cited claim to its cited chunk, since a claim that is true against some other chunk is still a mis-citation. Name the common failure in the prompt: any number, date or condition the evidence does not contain makes the claim unsupported.

The gate turns verdicts into a decision. Fully grounded ships as is. Mostly grounded drops the unsupported sentences and keeps the rest. Below a floor, refuse. The thresholds come from labeling a few hundred real answers and picking the cut that gives your domain the precision it needs, looser for password resets, strict for dosages and contract terms. Log every verdict, because silently redacted sentences hide your own failure rate, and half the time a failed claim points at a retrieval gap rather than a generation problem. Grounding confirms the answer matches the evidence, not that the evidence matches reality, so it does not replace retrieval hygiene. The extraction, verification and gate code is in stopping RAG hallucinations with a grounding gate.

Where to start

If you have none of this, do not build all of it. Build it in the order the failures arrive.

  1. Write twenty cases from the last month's incidents, each with must_include, must_not_include and a one-sentence rubric. Wire the scorer and the nonzero exit into CI as a required check. This is a weekend of work.
  2. Pin the judge model version and label a couple of hundred cases by hand. Measure kappa at your gate threshold. Below 0.6, tighten the rubric, switch model family, or route the disagreements to a human until it clears.
  3. Turn every production failure into a new eval case in the same pull request as the fix. This one habit stops the set from rotting and is cheaper than any pipeline.
  4. If you run an agent, add trajectory invariants to your recorded runs: must-call, must-not-call, a step budget. Then run each case ten to twenty times and gate on consistency and the Wilson floor instead of pass@1.
  5. If the agent holds a conversation, add a handful of adversarial simulator scenarios covering the money paths and the safety policies, scored on sandbox state and policy separately.
  6. If you serve RAG, add the grounding gate at the API boundary and build the labeled retrieval set you will need the first time someone wants a new embedding model.
  7. Put every model id behind a registry, and make the replay gate plus shadow and canary the only way a version changes.
  8. Measure coverage of the set against live traffic and add the health gate. Only then point an optimizer at the set, because an optimizer against a stale set or an uncalibrated judge makes the product worse with a better number.

The teams that ship LLM features without fear are not the ones with the best prompts or the newest model. They are the ones who can change a prompt, a model or an index on a Tuesday and know within minutes whether they broke something, because a set of real cases, scored by a grader they have checked and refreshed from the traffic it describes, stands between the change and their users. Observability tells you what went wrong. Evals stop it from going wrong, and every piece above is ordinary code.

Building this?I help teams put this into production. Thirty minutes, no pitch.