AI Agent Reliability in Production: What Breaks and What Fixes It

Every way an AI agent fails in production, from bad tool calls to blown error budgets, with the pattern that fixes each one and the essay that goes deep.

Guide15 essays20 min readUpdated

Most of the agent reliability work I do starts the same way. A team has an agent that works in the demo and holds up for a few weeks in production, and then the failures start arriving in a shape nobody planned for. Not crashes. A tool call retried into a dead endpoint. A customer charged twice. A run that called the same search seventy times. A report that was confidently wrong because of a decision made six steps before the error showed up. The model in every one of these cases was fine. The layer around the model had no idea what to do when something went wrong.

That is the thesis of this guide, and of the fifteen essays it pulls together: AI agent reliability is an engineering problem that lives in your code, not a model problem that a bigger model will solve for you. The failures are specific, they repeat across teams, and each has a pattern that fixes it.

I have grouped them the way they tend to come up: the tool boundary first, then retries, stopping, surviving a crash, load, the two tricks for a single step that will not behave, measurement, and debugging. Read it end to end, or jump to the section that matches the incident you had this week. Each one links the essay that goes into the code in full.

Most failures happen at the tool boundary

Watch an agent fail in production and it rarely looks like bad reasoning. The model picks the right tool and fills in plausible arguments, then the call comes back with a 400, a 401, or nothing at all, and the harness does the one thing you did not want: it tries the same arguments again.

The fix is a disciplined layer between the model and the tool that does three things. Validate arguments before you execute, so a missing field or an out of range number becomes a cheap self correction on the model's next turn, not a side effect. Classify every failure into a category, because a timeout and a bad argument look identical to a naive retry loop and need opposite responses. Then recover per category with a hard attempt cap, so transient failures get backoff, bad arguments go back to the model to repair, auth failures escalate once, and everything else surfaces as an honest failure instead of a silent hang.

from enum import Enum
import httpx

class Failure(str, Enum):
    TRANSIENT = "transient"          # retry with backoff
    BAD_ARGS = "bad_args"            # hand the error back to the model
    AUTH = "auth"                    # refresh or escalate, never loop
    UNRECOVERABLE = "unrecoverable"  # stop and report clearly

def classify(err: Exception) -> Failure:
    if isinstance(err, (httpx.TimeoutException, httpx.ConnectError)):
        return Failure.TRANSIENT
    if isinstance(err, httpx.HTTPStatusError):
        code = err.response.status_code
        if code in (408, 429, 502, 503, 504):
            return Failure.TRANSIENT
        if code in (400, 422):
            return Failure.BAD_ARGS
        if code in (401, 403):
            return Failure.AUTH
        if code >= 500:
            return Failure.TRANSIENT
    return Failure.UNRECOVERABLE

That single distinction, 429 and 5xx to transient while 400 and 422 go back to the model, is the difference between an agent that recovers and one that hammers a dead endpoint. I walked through the full validate, classify, recover loop in tool call reliability.

The other half of the boundary is what comes back. Once model output feeds code instead of a person, "mostly valid" JSON is a failed handoff. JSON mode guarantees the syntax parses and nothing else: the field can be misnamed, the number can arrive as a formatted string, the enum can be a value you never defined. Route every structured call through one gateway that owns the schema, parses in two gates (syntax, then schema), repairs with the exact validation error rather than a blind retry, bounds the repair, and raises loudly when it runs out. I built that gateway in both Pydantic and Zod in structured output reliability.

Retries that are safe to have

Retries make an agent tolerant of the network. They also create two failures of their own, both quieter than the problem they were meant to solve.

The first is the duplicate side effect. When a write tool's response is lost after the remote committed, the committed call and the never-arrived call are indistinguishable from where your code sits, so you retry, and sometimes you retry something that already happened.

The fix is not fewer retries. It is idempotency keys on every write tool, derived from the business action and never from a timestamp or a random value generated inside the retry loop. The same order and the same amount produce the same key on every attempt; a different order gets its own. Pass the key to the provider where it supports one, because that covers the commit-to-acknowledgement gap you cannot see, and enforce it yourself with a dedup store everywhere else. The store has to reserve the key before doing the work, not after, or two overlapping attempts both sail through the check. Tag tools as read or write and only wrap the writes. The key derivation, the dedup store, and the TTL and failure-branch decisions are in idempotency keys for retry-safe side effects.

The second retry failure hides in code that looks completely reasonable. A step fails, the orchestrator appends the error to the running message list and asks the model to try again. The failed attempt is still in the window, so every retry is conditioned on the mistake it is supposed to fix. The model cannot choose to ignore the part above where it got it wrong; the most probable continuation is something close to what it just did.

Retry on a clean context instead. Each attempt gets a fresh window with the original task and one distilled correction, a single sentence of at most 25 words produced by a small model from the failure, and the raw transcript goes to your logs. Bound the attempts, and escalate when two consecutive corrections say the same thing. Reserve this for failures where the model's own output was wrong; a timeout or a 503 should be retried as-is with backoff, because the context was never the problem. The full retry policy is in clean-context retries.

Runs that stop when they should

A crash at least stops billing you. The failures in this section do not. They keep the meter running while producing nothing, and you learn about them from a cost alert or a support ticket.

The first is the loop. A ReAct-style agent calls a search tool, gets an unhelpful result, decides one more identical call will help, and does it again. Step repetition is the single most common way agents fail in production, around one in six failures, and a hard step ceiling is the wrong instrument for it. The ceiling measures the quantity of steps; what you care about is whether the steps are going anywhere.

A no-progress guard fingerprints each tool interaction and watches for repeats.

import { createHash } from "node:crypto";

// Arguments AND a result digest go into the key, so a retry that returns
// something new, or a paginated call with a moving cursor, is not a repeat.
function fingerprint(name: string, args: unknown, result: unknown): string {
  const canonical = JSON.stringify({ name, args, result: digest(result) });
  return createHash("sha1").update(canonical).digest("hex");
}

class NoProgressGuard {
  private recent = new Map<string, number>();
  constructor(private limits: { default: number; perTool: Record<string, number> }) {}

  // True when this exact interaction has repeated past its threshold.
  record(name: string, args: unknown, result: unknown): boolean {
    const key = fingerprint(name, args, result);
    const seen = (this.recent.get(key) ?? 0) + 1;
    this.recent.set(key, seen);
    return seen >= (this.limits.perTool[name] ?? this.limits.default);
  }
}

Same tool, same arguments, same result, showing up again, is the definition of an agent that has learned nothing. Per-tool limits give pollers like check_job_status headroom without turning the guard off, and a second tracker that digests the run's working state catches the subtler stall where the agent varies its calls and still gets nowhere. When either fires, escalate rather than kill: a system message saying the path is dead breaks most loops on its own, removing the offending tool breaks most of the rest, and only then do you return a structured stopped_no_progress result with partial state. The full ladder is in the no-progress guard.

The second is the run nobody is waiting for. A user closes the tab after eight seconds, and on the server nothing notices. The agent finishes its model call, fires three tool calls, spawns a subagent, and forty seconds later serializes an answer for nobody. An outer timeout does not help, because it fires on the top of the stack and the leaves keep running. You have stopped listening, not stopped spending.

Two disciplines share one path here. Pass an absolute deadline, a single point on the monotonic clock computed once at the top and handed down unchanged, so ten careful steps cannot each spend the full budget the way per-step timeouts allow. And propagate cancellation down the tree: cancel the parent task and the children unwind, siblings of a failed call get cancelled rather than left to finish, and every layer catches the cancellation only to clean up and always re-raises it. Wire the client disconnect to cancelling the top-level task and the external cancel is nearly free. The asyncio mechanics are in deadline and cancellation propagation.

Runs that survive a crash

A sixty-step research run dies at step forty when a deploy rolls the process. When it comes back it starts from step one, pays for forty steps a second time, and re-fires every side effect along the way. A long agent run is not a function call, and treating it like one is what makes agents expensive and untrustworthy.

Durable execution changes the unit of recovery from the whole run to the single step. Write a durable record after each step completes, atomically, so the run becomes a log of finished steps and what each produced. On resume, read the log and skip past everything already done. The contract that makes this work is replay safety: every non-deterministic step, a model call, a search, a timestamp, records its result the first time and returns the recorded value on replay instead of executing again. Let those calls run live on resume and you have started a subtly different run, not recovered the old one. Step ids have to be stable, keyed on position, never on time or a random uuid, or the resumed run matches nothing and re-runs everything.

Recording handles reads. Writes have a narrow window the checkpoint cannot close: the effect fires, then the record lands, and a crash between the two replays the effect. This is where the idempotency keys from the previous section come back, derived from run and step so they are identical on every replay. At-least-once execution plus an idempotent effect gives you effectively exactly-once across a boundary you do not control. Build the small version first, and adopt an engine like Temporal or a LangGraph checkpointer when you need multiple workers, timers, and signals. The hand-rolled version and its tradeoffs are in durable execution for agents.

Load in both directions

An agent is a load generator that decides for itself how much load to create. Ask it to enrich four hundred accounts and a reasonable plan spawns a sub-task per account, each firing three or four tool calls, and the run wants fifteen hundred outbound calls right now against an API that tops out at ten a second. The only thing between you and that was a Promise.all.

Four quantities are unbounded when this melts down, and each needs its own control. The work queue needs a bounded queue that refuses work when full, because a queue with no maximum is a memory leak with good intentions. Concurrency needs a semaphore sized to the slowest thing you call. Request rate needs a token bucket, with the cost set to estimated tokens for providers that limit on tokens per minute. And the token budget needs a hierarchical ceiling, where a parent debits a child's grant up front so no branch can drain the pool the others need. Add a circuit breaker so a dependency that is already failing is skipped for a cooldown instead of tying up a slot on every doomed call. All four controls are in backpressure and rate limiting for agents.

The other direction is the provider failing you. Early this month three major providers degraded within an afternoon, and products that called one provider and waited went dark with it. Its p99 is your p99. Its outage is your outage. The trap is retrying your way through a real outage: every request burns its full retry budget waiting, the worker pool fills with sleeping requests, and you have converted one provider's outage into your own resource exhaustion. Retries are for blips. Failover is for outages. A circuit breaker per provider, ideally per model, tells the two apart and fails fast once a provider has clearly gone bad.

Wrap the breakers in a gateway that tries providers in order, skips open circuits in microseconds, and caps total attempts across all providers, because fallbacks are usually pricier and an uncapped failover is how a bad afternoon becomes a five-figure invoice. Hedge a small slice of traffic for the slow tail. Decide in advance what an honest degraded mode looks like when nothing is healthy, because a spinner that lies is the thing users remember. The gateway and the hedging code are in provider outage resilience.

A single step that will not behave

Sometimes the measurement points at one step: a classification that decides whether a ticket is auto-refunded or escalated, right nine times in ten on the same input. Fine-tuning is a project and the frontier model on every call blows the budget. There are two cheaper options, and both work from outside the model.

The first is to vote. Run the step several times in parallel at a nonzero temperature, canonicalize the answers so formatting differences do not split the tally, and take the majority. When the runs fail independently the arithmetic is lopsided in your favor: a step at 10% error drops to about 0.86% with five samples and 0.09% with nine. The trap is that wrong answers agree too. An ambiguous phrasing or a missing field pushes every sample the same wrong way, and five-of-five agreement on a correlated error looks identical from the outside to five-of-five on a correct one. So force the samples to be able to disagree, by varying the framing or drawing from more than one model, and split the outcome into three buckets: strong agreement auto-approves, middling agreement gets an independent check that is structurally different from the vote, and real disagreement escalates. Apply it to the one or two steps your measurements flagged, never as a blanket wrapper. The calculator and the early-stop sampler are in consensus sampling.

The second is to let the agent decline. Models guess because the whole training and evaluation stack rewarded guessing: a wrong answer and an abstention both score zero, so the expected value of a guess always wins. You will not prompt your way out of that. An abstention gate sits between the draft and the user. Compute a confidence signal you can get from outside, verbalized confidence as a field in the structured response and self-consistency disagreement across a few samples, then compare it to a threshold and choose to answer, abstain, or escalate. The threshold is the part teams skip. Calibrate it on a few hundred labeled questions so you can say "we answer 71% of questions and the ones we answer are wrong under 5% of the time", which is a service level a stakeholder can agree to. The gate and the calibration sweep are in the abstention gate.

Measuring reliability honestly

A 90% pass@1 eval score feels like a passing grade, and it lies in two ways at once.

The first lie is that an average hides how often you flip. pass@1 runs each case once, so it cannot tell an agent that passes 180 cases every time and fails 20 every time from one that passes every case 90% of the time. The first is a product. The second is a slot machine, and every customer pulls the lever once. The metric that captures this is pass^k, the probability the agent succeeds on all k attempts, the pessimist's version of pass@k. Run each case 10 to 20 times at production temperature and report the fraction of cases that passed every time next to the mean.

The second lie is that 90% steps do not make a 90% task. Per-step success multiplies.

from functools import reduce
from math import sqrt

def task_success(step_probs: list[float]) -> float:
    """End-to-end success of an independent-step pipeline."""
    return reduce(lambda acc, p: acc * p, step_probs, 1.0)

# 1 step  @ 0.90 -> 90.0%    5 steps @ 0.90 -> 59.0%
# 3 steps @ 0.90 -> 72.9%   10 steps @ 0.90 -> 34.9%

def wilson_lower_bound(passes: int, n: int, z: float = 1.96) -> float:
    """Honest floor on a success rate from n trials. 3 of 3 -> 0.44."""
    phat = passes / n
    denom = 1 + z * z / n
    center = phat + z * z / (2 * n)
    margin = z * sqrt((phat * (1 - phat) + z * z / (4 * n)) / n)
    return (center - margin) / denom

Instrument the trajectory and compute per-step reliability, because in a multiplicative system the weakest step dominates and that is where every hour of effort should go. Report a Wilson lower bound rather than the raw fraction so a lucky three-of-three cannot fool you, and gate releases on suite-level consistency and a per-case floor. The suite runner and the CI gate are in pass^k and per-step consistency.

Production needs the same honesty. A single 99.1% number on a dashboard, green, told one team they were fine while their most important customer waited eleven seconds for every correct answer. An agent fails along dimensions that do not correlate, correct but slow, fast but wrong, accurate but over budget, and averaging them lets a real regression hide behind strong numbers elsewhere. Define five or six SLOs, one per dimension that breaks, each as a good-event ratio over a window with an honest denominator. The gap to 100% is an error budget you are allowed to spend, and once you compute budget remaining per SLO the conversation with product changes from "the agent feels flaky" to "grounding has 12% of its monthly budget left, so the retrieval work jumps the queue". Alert on burn rate across a fast window and a slow window so you catch both the sudden outage and the quiet bleed, and make the release gate ask the budget instead of the room. The SLI computation and the alerting are in agent SLOs and error budgets.

Debugging a failure you cannot reproduce

A user sends a screenshot of your agent moving a meeting to the wrong day. You have good logs. You run the same prompt against staging ten times and it is right every time. The run that broke was one specific path through a system that takes a slightly different path on every execution, and the conditions are gone.

Nondeterminism enters from four places, and the model is only one. The model call varies even at temperature zero through batching and routing. Tools call the real world, and what a search returned at that instant is gone. The clock changes what a prompt with "today is" in it says. And the plumbing, random ids, parallel resolution order, retries that fired once and not again, shifts the path. You cannot make any of it deterministic. What you can do is record what each source produced, at a single interception point per boundary, onto an append-only tape, and in replay mode serve the recording without calling anything live. Key events by identity, the tool call id or the step index, never by position, or parallel calls break the replay. Record broadly, because the failure a customer complains about usually returned a 200. This is the opposite direction from checkpoint recovery: recovery continues forward with live calls, replay refuses them. The tape, and the environment that routes the clock and random generator through it, are in deterministic replay.

Once you can replay, you can find the step that actually broke the run, which is almost never the step that threw. A retrieval step at step three pulls the wrong document and returns confidently. Everything after it is competent work on a bad premise, and step nine is just where the premise finally collides with something that checks it. The thing you want is the decisive step: the earliest step whose correction flips the outcome from failure to success. Find it with counterfactual replay. Pin the steps before a candidate to their recorded outputs, replace the candidate's output with what it should have produced given only the context it actually had, replay the rest, and check whether the run now succeeds. Binary search the trajectory and a dozen replays become three or four; rank candidates with a cheap critic first so the first replay lands near the answer. The replay primitive and the search are in failure attribution with counterfactual replay.

Where to start

If you have none of this, do not try to build all of it at once. Most of these patterns earn their place only after a specific failure has shown up, and the order below is roughly the order the failures tend to arrive.

  1. Put a classifier and an attempt cap on every tool call, and a schema gateway on every structured output. This is the cheapest work on the list and it removes the largest share of production failures.
  2. Add idempotency keys to every write tool. Do this before you make retries more aggressive, not after.
  3. Switch retries of reasoning failures to clean contexts. One distilled sentence, a fresh window, a bounded attempt count.
  4. Add a no-progress guard and an absolute deadline with cancellation. These two stop the runs that bill you for nothing. Keep the step ceiling as a backstop behind the guard.
  5. Checkpoint long runs. Only once you have runs long enough and expensive enough to hurt when they restart.
  6. Bound the fan-out and put a breaker and a failover gateway in front of the provider. The signal is the first rate-limit storm, the first surprise invoice, or the first provider outage that became yours.
  7. Change what you measure. Run the eval suite k times, gate releases on consistency, and split the production dashboard into SLOs with error budgets. This is the step that tells you which of the earlier ones you still need.
  8. Vote or abstain on the step the measurement flags. Consensus for a discrete decision that flips, an abstention gate for an answer the agent should sometimes decline to give.
  9. Record every run so you can replay it, then attribute failures to the decisive step. Every failure becomes one you can reproduce, and every fix lands on the cause.

None of this requires a smarter model. It requires treating the layer around the model as the production system it already is, with the same instincts backend engineers have used for years: validate at the boundary, make retries safe, bound every loop, checkpoint long work, limit load, measure what customers feel, and reproduce before you fix. The model will keep getting better on its own. The code around it is yours.

Building this?I help teams put this into production. Thirty minutes, no pitch.