Your Evals Are Green and the Product Is Getting Worse
A frozen eval suite stops describing your product within weeks. New intents, new tools, new retrieval indexes and new traffic shapes all land after the set was written, so the build stays green while real users hit failures the set has never seen. This is how to measure eval-set drift against live traffic, mine novel production failures, promote them into a living golden set, and gate releases on coverage and freshness instead of a pass rate alone.
A team I worked with had a clean eval suite. Two hundred cases, ran in CI on every pull request, a green check next to a pass rate that sat around 94% for months. They trusted it, which was the whole point of building it. Then support tickets started climbing, slowly at first and then not slowly, and the eval suite stayed green through all of it. Nothing had regressed on the dashboard. Everything had regressed in the product.
What happened was not a model change and not a bug in the harness. The eval set had been written in the spring against the product as it existed in the spring. Over the summer the team shipped three new intents, swapped the retrieval index twice, rewrote half the system prompt, and added a batch of tools. The eval set did not know about any of it. It kept testing a product that no longer existed, scored it 94%, and said ship. Meanwhile the actual traffic had moved somewhere the set had never looked.
This is eval-set drift, and it is the failure mode that undoes most of the value of having evals in the first place. You build the suite to earn trust, and then the suite quietly stops deserving it while still showing you the number that earned it.
Why a frozen set ages off your product
A traditional test suite can sit still because the thing it tests sits still. add(2, 2) returns 4 this year and next year. An LLM product is not like that. The thing under test is a moving target made of prompts, tool schemas, retrieval corpora and a model, and the thing it serves, your traffic, moves on its own schedule for reasons you do not control.
Every release ages the set a little. A new feature adds intents the set has no cases for. A prompt rewrite changes which phrasings trip the system up, so your old adversarial cases test a weakness that no longer exists while the new weakness goes unmeasured. A reindex changes what retrieval returns, so cases that used to exercise grounding now pass trivially. And underneath all of that, users keep inventing inputs nobody designed for.
The result is two kinds of rot running at once. Coverage decay: the set stops containing the kinds of inputs your users actually send. And relevance decay: the cases it does contain test behavior that no longer matters. Both push the pass rate up and the real quality down, because the easy, stale cases keep passing while the hard, current cases are simply absent.
The CI eval post I wrote earlier names this in one line, "the dataset rots," and leaves it there. This is the full treatment, because naming it is not enough. If your set is frozen, the only question is how many weeks until it lies to you.
Measure the drift before you trust the number
You cannot manage drift you cannot see, so the first job is to make the gap between your eval set and your traffic into a number. The cheap, effective version: embed a sample of recent production inputs and your eval-set inputs in the same space, cluster the production side, and check how many of those clusters have an eval case anywhere near them. Clusters with no nearby case are the parts of your product that ship untested.
import numpy as np
from sklearn.cluster import KMeans
def coverage_report(prod_embeddings, eval_embeddings, n_clusters=40,
covered_threshold=0.80):
"""Fraction of production clusters that have at least one eval case nearby.
Embeddings are L2-normalized, so a dot product is cosine similarity."""
km = KMeans(n_clusters=n_clusters, n_init="auto", random_state=0)
labels = km.fit_predict(prod_embeddings)
centroids = km.cluster_centers_
centroids /= np.linalg.norm(centroids, axis=1, keepdims=True)
# Best cosine similarity from each cluster centroid to any eval case.
sims = centroids @ eval_embeddings.T # (n_clusters, n_eval)
nearest = sims.max(axis=1) # best match per cluster
covered = nearest >= covered_threshold
sizes = np.bincount(labels, minlength=n_clusters)
# Weight coverage by how much traffic each cluster actually carries.
traffic_coverage = sizes[covered].sum() / sizes.sum()
return {
"cluster_coverage": float(covered.mean()),
"traffic_coverage": float(traffic_coverage),
"uncovered_clusters": np.where(~covered)[0].tolist(),
}
The number that matters is traffic_coverage, not cluster_coverage. A set can cover 38 of 40 clusters and still miss a third of your traffic if the two uncovered clusters are where the volume is. Watch it over time. The day it starts falling is the day your green checkmark started drifting away from reality, and you will see it weeks before the support tickets do.
Mine production for the cases you are missing
Coverage tells you that you have gaps. Filling them means pulling real traces out of production and turning the useful ones into eval cases. The useful ones are of two kinds: traces that failed, and traces that are unlike anything your set already contains. You want both, because failures tell you where you are weak and novelty tells you where you are blind.
You rarely have ground-truth labels on live traffic, so lean on the weak signals you already emit. An explicit thumbs-down, an online LLM judge score below a bar, a tool error, a fallback to a cheaper model, a user who rephrased the same question three times, a conversation that escalated to a human. None of these is proof of failure on its own. Together they are a good enough filter to surface candidates worth a human glance.
from dataclasses import dataclass
@dataclass
class Trace:
id: str
text: str # the user input (PII already scrubbed upstream)
embedding: np.ndarray
judge_score: float | None # online judge, 0..1, if scored
thumbs_down: bool
tool_errors: int
retries: int
escalated: bool
def failure_signal(t: Trace) -> float:
"""Cheap 0..1 suspicion score. Not truth, just a ranking to triage by."""
score = 0.0
if t.judge_score is not None and t.judge_score < 0.6:
score += 0.4 * (0.6 - t.judge_score) / 0.6
if t.thumbs_down: score += 0.4
if t.tool_errors: score += 0.2
if t.retries >= 2: score += 0.2
if t.escalated: score += 0.3
return min(score, 1.0)
def novelty(t: Trace, eval_embeddings: np.ndarray) -> float:
"""1 - best similarity to any existing eval case. High means unfamiliar."""
return float(1.0 - (eval_embeddings @ t.embedding).max())
def candidates(traces, eval_embeddings, suspicion=0.3, min_novelty=0.25):
for t in traces:
if failure_signal(t) >= suspicion or novelty(t, eval_embeddings) >= min_novelty:
yield t
Run this over a window of traffic and you get a candidate pool: the traces worth a human deciding whether they belong in the set. The point is not to automate the decision. It is to cut the thousands of daily traces down to the few dozen that carry new information, so a person spends their attention where it pays.
Promote patterns, not noise
The trap at this stage is adding every failing trace directly to the set. Do that and you drown the set in near-duplicates of whatever broke this week, and you teach it to overfit to yesterday's incident. The fix is to promote representatives of clusters rather than individual traces, so each distinct failure pattern contributes a case or two instead of forty.
def select_for_promotion(cands, max_per_cluster=2, n_clusters=15):
"""Cluster the candidate pool and return the most suspicious few per
cluster. Turns a noisy pile into a short list of distinct patterns."""
cands = list(cands)
if len(cands) <= n_clusters:
return cands
X = np.vstack([c.embedding for c in cands])
labels = KMeans(n_clusters=n_clusters, n_init="auto",
random_state=0).fit_predict(X)
chosen = []
for cluster_id in set(labels):
members = [c for c, l in zip(cands, labels) if l == cluster_id]
members.sort(key=failure_signal, reverse=True)
chosen.extend(members[:max_per_cluster])
return chosen
What comes out of this is a short list, maybe twenty traces standing in for twenty patterns. Those go to a human to confirm the failure is real and to write the expected behavior, because a case with no trustworthy expected output is worse than no case at all. The discipline that makes this stick: a production failure becomes a new eval case in the same pull request that fixes the underlying bug. The case fails before the fix and passes after, which proves the fix worked and permanently vaccinates the set against that pattern coming back.
Gate on coverage and freshness, not just the pass rate
Mining is only half the loop. The other half is refusing to trust a set that has gone stale, which means the pass rate stops being the only thing standing between a change and production. Add two more gates: coverage of recent traffic, and the age of the set relative to how fast you ship.
import time
def eval_set_health_gate(coverage, median_case_age_days, releases_since_added,
min_traffic_coverage=0.75, max_median_age_days=30):
"""Returns (ok, reasons). A green pass rate on a stale set is not a pass."""
reasons = []
if coverage["traffic_coverage"] < min_traffic_coverage:
reasons.append(
f"traffic coverage {coverage['traffic_coverage']:.0%} "
f"below floor {min_traffic_coverage:.0%}; "
f"{len(coverage['uncovered_clusters'])} clusters untested"
)
if median_case_age_days > max_median_age_days:
reasons.append(
f"median case age {median_case_age_days}d exceeds "
f"{max_median_age_days}d; set is aging off the product"
)
if releases_since_added > 10:
reasons.append(
f"{releases_since_added} releases since a case was added; "
f"intake pipeline looks stalled"
)
return (len(reasons) == 0, reasons)
Wire this into the same CI job that runs the suite. A build can now fail not because a case regressed but because the set no longer describes the product well enough for its pass rate to mean anything. That reframes the whole thing. The pass rate answers "did we break a known case," and the health gate answers "do we still know enough cases for that to matter." You need both, and most teams ship with only the first.
The parts that bite
A few things go wrong in practice, and they are worth building against from the start.
Feedback loops poison the set. If your intake relies only on an LLM judge and you promote whatever that judge scores low, you encode the judge's blind spots into the very set you use to trust the judge. Keep a human in the promotion decision, and keep a slice of cases that were labeled by hand and never touched by the judge, so you have an anchor the loop cannot corrupt.
PII rides along in real traffic. Production traces contain real user data, and an eval set is a dataset you copy, share with contractors and paste into prompts. Scrub before anything lands in the set, not after, and keep the raw traces in a separate store with tighter access.
The set grows without bound. Every week adds cases and nobody ever removes any, so your suite takes an hour to run and still overweights a bug from last March. Cap contributions per cluster, retire cases that have passed untouched for many releases and no longer match live traffic, and keep a small stable core of canonical happy-path cases that you deliberately never churn. A living set means things leave, not only arrive.
Novelty is not failure. A trace can be unlike your set and still be a case the system handles fine. Novelty earns a trace a human glance, not an automatic promotion. Promoting every unfamiliar input teaches the set to chase the long tail and starve the common paths that actually carry your revenue.
The takeaway
An eval suite is not an artifact you build once and trust forever. It is a claim about your product, and the product moves out from under the claim a little with every release. Freeze the set and the claim goes stale in weeks while the dashboard keeps reporting the number that stopped being true. The fix is a loop: measure coverage against live traffic, mine the traces that failed or that you have never seen, promote clustered patterns rather than noise, and gate releases on freshness as hard as you gate on the pass rate. None of it is exotic. It is embeddings you already compute, signals you already emit, and a CI job you already run, wired into a habit of feeding the set from the same reality it is meant to measure.
If your evals have been green for a suspiciously long time while support volume creeps up, that gap is worth a hard look before the next release rides on it. Book a consultation call and we can pressure-test your eval set against your actual traffic and build the intake loop that keeps it honest.
Viral Ruparel
Generative AI consultant helping teams ship reliable LLM and agent systems in production.
Contact Viral about your AI project →Frequently Asked Questions
What is eval-set drift?+
Eval-set drift is what happens when your offline evaluation set stops representing the product it is supposed to measure. The set was written against a snapshot of your intents, prompts, tool schemas, retrieval indexes and traffic mix, and every release changes some of those. Within a few weeks the frozen set describes a product that no longer exists, so a build can pass every case while real users hit failures the set has never seen. It is distinct from model drift: the model can be perfectly stable and your eval set still drifts because the world around it moved.
How is eval-set drift different from eval drift or model drift?+
People use "eval drift" loosely for two different things. One is the evaluator drifting, for example an LLM judge whose scores creep as you tweak its prompt or swap its model. The other, which this post is about, is the eval set itself drifting away from production traffic. Model drift is a third thing: the system under test degrading over time. All three produce the same symptom, a green dashboard over a worse product, so the fix starts with naming which one you have. Measuring coverage of your eval set against recent traffic tells you whether the set is the problem.
How do I keep an eval set from going stale?+
Treat it as a living dataset with an intake pipeline, not a file you wrote once. Continuously sample production traffic, flag the traces that look like failures or like inputs your set has never covered, cluster those candidates so you promote patterns rather than one-off noise, label the representatives, and add them to the golden set in the same change that fixes the underlying bug. Then measure the set's coverage of recent traffic and its median age, and gate releases when either crosses a threshold.
Will promoting production traces into my eval set cause overfitting?+
It can, if you only ever add past failures and never prune. The set starts to overweight historical bugs and underweight the common happy paths that pay the bills. Guard against it by promoting clustered representatives rather than every failing trace, capping how many cases a single cluster contributes, keeping a stable core of canonical happy-path cases that you never churn, and retiring cases that have not failed in many releases and no longer reflect live traffic.
Related Articles
Your Agent's One Reliability Number Is Hiding the Failures That Matter
A single pass rate on a dashboard tells you an agent is "99% reliable" and tells you nothing about which 1% is on fire. Meanwhile the slow answers, the blown budgets, and the confidently wrong responses all average into the same green number. This is how to define real SLOs for an agent across the dimensions that actually break, track an error budget against each one, alert on burn rate before users feel it, and gate releases on the budget instead of on a vibe.
When Every AI Provider Goes Down at the Same Time
Early this month three of the big model providers degraded within a few hours of each other, and a lot of AI features went dark with them. If your product calls one provider and waits, their outage is now your outage. This is how to build a resilience layer in front of your model calls: a circuit breaker that fails fast, health-aware failover across providers, request hedging for the slow tail, and a degraded mode that keeps the lights on when nothing is healthy.
Your Agent Got the Right Answer the Wrong Way
Your eval checks the final answer and goes green. Meanwhile the agent called six tools to do the work of two, hit a write endpoint it never needed, and landed on the right output by luck. Output-only evals cannot see any of that, and the newer reasoning models take longer, more autonomous paths where it matters more. Trajectory evaluation grades the steps: which tools ran, in what order, with what arguments. Here is how to build it and gate it in CI.