Your Agent's One Reliability Number Is Hiding the Failures That Matter
A single pass rate on a dashboard tells you an agent is "99% reliable" and tells you nothing about which 1% is on fire. Meanwhile the slow answers, the blown budgets, and the confidently wrong responses all average into the same green number. This is how to define real SLOs for an agent across the dimensions that actually break, track an error budget against each one, alert on burn rate before users feel it, and gate releases on the budget instead of on a vibe.
Last month a team showed me their agent reliability dashboard with some pride. One big number, top left, 99.1%, green. They shipped on that number. Product asked "are we good?" and someone glanced at the dashboard and said yes, and a release went out.
The problem was that their most important customer, a design partner paying real money, was having a terrible week. The agent was answering their questions correctly and taking eleven seconds to do it, because that customer's data lived in a slower backend and the agent made four extra tool calls to assemble it. Correct, so it counted as a success. Slow enough that the customer was drafting an angry email. The 99.1% had quietly folded latency, cost, and correctness into one average, and the one dimension that was actually on fire averaged away to nothing.
That is the core problem with a single reliability score for an agent. It is not that the number is wrong. It is that it answers a question nobody should be asking. "Is the agent reliable" has no single answer, because an agent fails in several independent ways, and a blend across them is a number that can only mislead you.
One number is the wrong altitude
A plain web service mostly fails one way: it returns an error or it does not. One SLO on the error rate covers most of it. An agent is not that. It can be:
- Correct but slow. The answer is right and it took fourteen seconds.
- Fast but wrong. It returned in a second with a confident, incorrect answer.
- Correct and fast but expensive. It burned four times the tokens and three extra tool calls to get there.
- Accurate on average but ungrounded on a specific path. It cites real sources ninety percent of the time and invents them for one category of question.
These do not move together. A prompt change can cut latency and quietly wreck grounding. A model upgrade can improve correctness and double cost. When you average them, a real regression in one dimension hides behind strong numbers in the others, and the blend stays green while a specific failure mode ruins a specific customer's week.
The fix is not a better single metric. It is to stop asking for one. You define a small set of Service Level Objectives, one per dimension that actually breaks, and you track each on its own with its own budget. This is the "six SLOs, not one score" idea that has been going around the reliability circles this autumn, and it is exactly how SRE teams have run plain services for years. Agents just make the case for it undeniable, because they have more ways to fail.
From traces to SLIs
An SLO is the target. The thing you actually measure is the SLI, the Service Level Indicator, and you already have the raw material if you have any tracing on your agent. If you do not, that is the first job, and the OpenTelemetry GenAI conventions are the right place to start so your spans carry latency, token counts, tool outcomes, and a task result.
Assume each completed agent task boils down to a record like this. The point is to turn fuzzy behavior into plain fields you can count.
from dataclasses import dataclass
@dataclass
class TaskRecord:
task_id: str
succeeded: bool # did the task meet its success criterion
latency_s: float # end-to-end, user request to final answer
cost_usd: float # total model + tool cost for the task
grounded: bool # did grounding/faithfulness check pass (if applicable)
tool_errors: int # failed tool calls during the task
ts: float # unix timestamp the task completed
Each SLI is just a ratio of good events to valid events over a window. The trick is defining "good" per dimension and being honest about the denominator. A latency SLI only counts tasks that actually ran, not ones that errored out before they could be slow.
def compute_slis(records: list[TaskRecord]) -> dict[str, float]:
"""Return the good-event ratio for each SLI over the given records.
Each value is in [0, 1]; higher is better."""
n = len(records)
if n == 0:
return {}
success = sum(1 for r in records if r.succeeded) / n
fast = sum(1 for r in records if r.latency_s <= 8.0) / n
in_budget = sum(1 for r in records if r.cost_usd <= 0.05) / n
clean_tools = sum(1 for r in records if r.tool_errors == 0) / n
# Grounding only applies to tasks that retrieved, so scope the denominator.
grounded_records = [r for r in records if r.grounded is not None]
grounded = (
sum(1 for r in grounded_records if r.grounded) / len(grounded_records)
if grounded_records else 1.0
)
return {
"success": success,
"latency": fast,
"cost": in_budget,
"grounding": grounded,
"tool_health": clean_tools,
}
Notice that latency and cost become ratios too: the fraction of tasks that came in under a threshold. That is deliberate. "p95 under 8 seconds" is really "95% of tasks are under 8 seconds," which is the same shape as a success rate, so you can treat every dimension with one mental model and one error-budget calculation. Pick the thresholds from what your users actually feel, not from what the current system happens to produce.
Error budgets make the target spendable
Now attach an SLO target to each SLI. The gap between a perfect 100% and the target is the error budget, and the useful reframe is that the budget is a thing you are allowed to spend. A 99% success SLO over 28 days does not mean "never fail." It means you have a budget of 1% of tasks to fail, and as long as you are inside it, you are meeting your promise and you are free to ship features. Blow past it and the rule is simple: you stop shipping features and spend your effort on reliability until the budget recovers.
SLOS = {
"success": 0.98,
"latency": 0.95,
"cost": 0.90,
"grounding": 0.99,
"tool_health": 0.97,
}
def budget_status(slis: dict[str, float], slos: dict[str, float]) -> dict[str, dict]:
"""For each SLO, report how much error budget is left as a fraction.
1.0 means the full budget is intact; 0.0 means it is exactly spent;
negative means the SLO is being violated right now."""
out = {}
for name, target in slos.items():
sli = slis.get(name, 1.0)
budget = 1.0 - target # total allowed failure fraction
consumed = max(0.0, 1.0 - sli) # observed failure fraction
remaining = 1.0 - (consumed / budget) if budget > 0 else 1.0
out[name] = {
"sli": round(sli, 4),
"slo": target,
"budget_remaining": round(remaining, 3),
"healthy": sli >= target,
}
return out
The moment you compute budget remaining per SLO, the conversation with product changes. Instead of "the agent feels flaky," you can say "grounding has 12% of its monthly budget left and it is the eighth, so the retrieval work jumps the queue." A grounding regression and a cost regression now argue for their own fixes on their own timelines, which is the whole point of splitting the score. This pairs naturally with offline work like pass^k consistency checks: the eval tells you whether a change is good before release, the SLO tells you whether production is still good after.
Alert on burn rate, not on a tripped threshold
A budget you only check at month end is a budget you will always discover is gone too late. You want to know when you are spending it too fast to survive the window. That rate is the burn rate: how quickly you are consuming the budget relative to the pace that would use it up exactly on schedule. A burn rate of 1 means you finish the window having spent precisely the budget. A burn rate of 14 means you will be out in a fraction of the window if nothing changes.
The mistake is alerting on a single window. A short window catches a sudden outage but screams at every brief blip. A long window is calm but sleeps through a slow, steady regression until the budget is gone. So you run both at once and only page when a fast window and a slower window both agree.
def burn_rate(records: list[TaskRecord], sli_name: str,
slo: float, window_s: float, now: float) -> float:
"""Error budget burn rate over a trailing window.
1.0 = on pace to exactly exhaust the budget across that window."""
window = [r for r in records if now - r.ts <= window_s]
if not window:
return 0.0
slis = compute_slis(window)
observed_fail = max(0.0, 1.0 - slis.get(sli_name, 1.0))
budget = 1.0 - slo
return observed_fail / budget if budget > 0 else 0.0
def should_page(records, sli_name, slo, now) -> bool:
"""Page only when a fast and a slow window both breach their thresholds.
Fast window catches acute outages; slow window catches quiet bleed."""
fast = burn_rate(records, sli_name, slo, window_s=3600, now=now) # 1h
slow = burn_rate(records, sli_name, slo, window_s=6 * 3600, now=now) # 6h
return fast > 14.4 and slow > 6.0
Those thresholds are the standard multi-window numbers from the SRE playbook, and they are a starting point, not scripture. A 14.4 burn over an hour means you would spend a month of budget in two hours, which is worth waking someone for. A 6x burn sustained over six hours is the quiet killer that a one-hour window alone would shrug off. Tune the windows to your traffic volume: if you only see a few hundred tasks an hour, widen the fast window so a handful of failures does not look like an apocalypse.
Gate releases on the budget, not on a feeling
The last piece is where SLOs earn their keep: they turn "are we good to ship" from a hallway opinion into a check. Before a release promotes, ask the budget. If any SLO is already in violation, or a pre-release eval run burned a big slice of a budget, the release waits. This is the same instinct as gating on trajectory and output evals, extended to live reliability.
def release_gate(records: list[TaskRecord], now: float) -> tuple[bool, list[str]]:
"""Block the release if any SLO is unhealthy or burning fast right now.
Returns (allowed, reasons)."""
recent = [r for r in records if now - r.ts <= 24 * 3600]
status = budget_status(compute_slis(recent), SLOS)
blockers = []
for name, s in status.items():
if not s["healthy"]:
blockers.append(f"{name}: SLI {s['sli']} below SLO {s['slo']}")
elif s["budget_remaining"] < 0.25:
blockers.append(f"{name}: only {int(s['budget_remaining']*100)}% budget left")
return (len(blockers) == 0, blockers)
allowed, reasons = release_gate(records, now=time.time())
if not allowed:
raise SystemExit("Release blocked:\n " + "\n ".join(reasons))
The first time this gate blocks a release that everyone "felt" was fine, you will get an argument. That argument is the feature working. Someone was about to ship on a vibe, and the budget said the grounding SLO was already underwater. Shipping more features onto a reliability problem is how you turn a bad week into a bad quarter.
The traps worth knowing before you build this
Vanity SLOs are the first one. If you set every target to what the system does today, every SLO is permanently green and you have built an expensive thermometer that only reads "fine." Set targets from what users need, then let the gap between that and reality be the work.
Bad denominators quietly lie. If your latency SLI counts tasks that errored before they could be slow, a spike in errors will make latency look better, which is backwards. Scope each SLI to the events where the dimension even applies, the way the grounding SLI above ignores tasks that never retrieved.
Too many SLOs is as useless as one. If you track fifteen dimensions, nobody owns any of them and every incident has a plausible scapegoat. Five or six, each with a named owner and a clear remediation, beats a wall of dials nobody reads.
And correctness is the hard SLI, because "succeeded" is doing a lot of work in that record. For some agents it is a deterministic check against a known answer. For open-ended ones it is an LLM judge or a sampled human review, and both have their own error bars. Be honest that your success SLI is an estimate, sample enough to trust it, and do not let a noisy judge become the reason you page someone at 3am.
The takeaway
A single reliability score is comfortable precisely because it refuses to tell you anything actionable. It cannot say which failure mode is costing you a customer, because it averaged that failure mode into oblivion before it reached the dashboard. Split it. Pick the handful of dimensions where your agent actually breaks, give each an SLO and an error budget, alert on burn rate across a fast and a slow window, and make your release gate ask the budget instead of the room. It is a few hundred lines on top of tracing you should already have, and it changes reliability from a feeling into a number you can defend. The same split also makes you honest about the thing you cannot fully control, the providers underneath you, which is its own discipline worth having a resilience layer for.
If your agent reliability still lives in one green number and you want to break it into SLOs you can actually act on, book a consultation call and we can map your real failure modes and put budgets and gates around them.
Viral Ruparel
Generative AI consultant helping teams ship reliable LLM and agent systems in production.
Contact Viral about your AI project →Frequently Asked Questions
What is an SLO for an AI agent?+
An SLO (service level objective) for an agent is a target you commit to on one measurable dimension of its behavior over a window, for example "95% of support tasks resolve correctly over 28 days" or "p95 end-to-end latency stays under 8 seconds." The measurement itself is the SLI (service level indicator), computed from your traces. The gap between 100% and your SLO is the error budget: the amount of failure you have agreed is acceptable before you stop shipping features and fix reliability instead. Agents need several SLOs at once because they fail in several independent ways, so a single blended score hides the dimension that is actually breaking.
Why is one reliability score not enough for an agent?+
Because an agent fails along dimensions that do not correlate. It can be correct but slow, fast but wrong, accurate but three times over budget, or grounded on average while hallucinating on a specific tool path. Averaging those into one number lets a serious regression in one dimension hide behind strong results in the others. You ship on green and your highest value customers hit the failure the blend smoothed over. Separate SLOs per dimension keep each failure mode visible and individually accountable.
What dimensions should an agent SLO cover?+
Start with the handful that map to money and trust: task success or resolution rate, end-to-end latency (p95 and p99, not the mean), cost per successful task, grounding or faithfulness pass rate for anything that retrieves, and tool-call error rate. Add a safety or policy-violation SLO if the agent takes actions with side effects. Keep the set small enough that every SLO has an owner and a clear remediation when its budget burns.
How does burn-rate alerting work for agents?+
Burn rate is how fast you are spending the error budget relative to the rate that would exhaust it exactly at the end of the window. A burn rate of 1 means you will use the whole budget and no more. You alert on multiple windows at once: a fast window (an hour or so) with a high threshold catches a sudden outage, and a slow window (six hours or a day) with a lower threshold catches a quiet, steady regression that a short window would miss. This is the same multi-window approach the Google SRE workbook uses, applied to agent SLIs instead of HTTP error rates.
Related Articles
When Every AI Provider Goes Down at the Same Time
Early this month three of the big model providers degraded within a few hours of each other, and a lot of AI features went dark with them. If your product calls one provider and waits, their outage is now your outage. This is how to build a resilience layer in front of your model calls: a circuit breaker that fails fast, health-aware failover across providers, request hedging for the slow tail, and a degraded mode that keeps the lights on when nothing is healthy.
Your Agent Got the Right Answer the Wrong Way
Your eval checks the final answer and goes green. Meanwhile the agent called six tools to do the work of two, hit a write endpoint it never needed, and landed on the right output by luck. Output-only evals cannot see any of that, and the newer reasoning models take longer, more autonomous paths where it matters more. Trajectory evaluation grades the steps: which tools ran, in what order, with what arguments. Here is how to build it and gate it in CI.
Your Agent Would Rather Guess Than Admit It Doesn't Know
Most agents never say "I don't know." They were trained on benchmarks that reward a confident guess over an honest abstention, so in production they invent a policy, a number, or a citation and say it with a straight face. An abstention gate fixes that: a cheap confidence signal, a threshold you calibrate against a real error budget, and a decision to answer, defer, or escalate. This is how you get an agent that knows its own limits.