Your Agent Would Rather Guess Than Admit It Doesn't Know
Most agents never say "I don't know." They were trained on benchmarks that reward a confident guess over an honest abstention, so in production they invent a policy, a number, or a citation and say it with a straight face. An abstention gate fixes that: a cheap confidence signal, a threshold you calibrate against a real error budget, and a decision to answer, defer, or escalate. This is how you get an agent that knows its own limits.
A support agent I reviewed last quarter told a customer their enterprise plan included a feature it had never shipped. The customer quoted it back to their account manager. The account manager escalated it as a broken promise. Nobody had lied on purpose. The model was asked a question it did not have grounding for, and instead of saying so, it produced a fluent, plausible, completely invented answer and delivered it with total confidence.
That is the failure mode that quietly costs the most. Not the agent that crashes, which you notice, but the agent that is wrong and sounds right, which you ship. And the reason it happens is not a bug you can patch. It is baked into how these models were trained and measured.
Why your model guesses on purpose
Think about how a model gets scored during training and evaluation. A benchmark asks a question. A correct answer scores one. A wrong answer scores zero. An abstention, an "I don't know," also scores zero. Now do the expected-value math the way the optimizer does. If you guess, you have some chance of scoring one and otherwise score zero. If you abstain, you score zero for certain. Guessing weakly dominates. The model that always produces an answer beats the model that sometimes declines, every time, on the leaderboard.
RLHF makes it worse rather than better. Human raters, given a long confident answer and a short "I am not sure," tend to prefer the confident one. So the reward model learns that confidence reads as quality, and the policy learns to sound sure even when it is not. By the time a frontier model reaches you, it has spent its entire life being rewarded for guessing and punished for humility. Then you put it in front of a customer and act surprised when it makes things up.
You are not going to retrain that behavior out with a better prompt. "Only answer if you are sure" helps a little and fails often, because the model's internal sense of "sure" is exactly the thing that got miscalibrated. What you can do is stop treating the model's first answer as the final answer. Put a gate in front of it.
The abstention gate
The idea is small and it borrows from a decades-old technique called selective prediction. Instead of a system that always answers, you build one that can decline. Every response passes through three steps. You produce a draft answer. You compute a confidence signal for that draft. Then a policy compares the signal against a threshold and picks one of three actions: return the answer, abstain and tell the user the agent does not know, or escalate to a stronger model or a human.
The whole design lives in two questions. What signal tells you the answer is shaky, and where do you set the line? Get those right and the rest is plumbing.
A confidence signal you can actually get
You usually cannot see the model's token probabilities through a production API, and even when you can, raw logits are badly calibrated for this. So use signals that work from the outside.
The cheapest is verbalized confidence. Ask the model to rate its own certainty as part of a structured response. It is not perfectly calibrated on its own, but it is stable under paraphrasing and it is nearly free.
import json
from openai import OpenAI
client = OpenAI()
ABSTAIN_SYSTEM = """You are a careful assistant. Answer only from the context
you are given. Respond as JSON with three fields:
answer: your best answer, or null if the context does not support one
confidence: your calibrated probability the answer is correct, 0.0 to 1.0
reason: one short sentence on what the confidence is based on
Do not inflate confidence. A wrong confident answer is worse than "I don't know"."""
def draft_with_confidence(question: str, context: str) -> dict:
resp = client.chat.completions.create(
model="claude-sonnet-5-5", # your normal workhorse
messages=[
{"role": "system", "content": ABSTAIN_SYSTEM},
{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
],
response_format={"type": "json_object"},
temperature=0,
)
return json.loads(resp.choices[0].message.content)
The stronger signal, and the one I trust more, is self-consistency disagreement. Ask the same question a handful of times at a non-zero temperature and see how much the answers move. If the model lands on the same answer every time, that agreement is meaningful. If it says four different things across four samples, it does not know, no matter how confident any single response sounded. This signal needs no access to internals and no training, and it catches the case verbalized confidence misses, where the model is fluently and consistently wrong versus fluently and randomly wrong.
from collections import Counter
def consistency_score(question: str, context: str, n: int = 5) -> float:
"""Fraction of samples that agree with the most common answer.
1.0 means every sample said the same thing; low means the model is guessing."""
answers = []
for _ in range(n):
resp = client.chat.completions.create(
model="claude-sonnet-5-5",
messages=[
{"role": "system", "content": ABSTAIN_SYSTEM},
{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
],
response_format={"type": "json_object"},
temperature=0.7, # need spread to see disagreement
)
ans = json.loads(resp.choices[0].message.content).get("answer")
# normalize so trivial wording differences don't read as disagreement
answers.append(str(ans).strip().lower() if ans is not None else "__abstain__")
top, count = Counter(answers).most_common(1)[0]
return count / len(answers)
Combine them however your data says to. A simple weighted blend of verbalized confidence and consistency works surprisingly well, and it degrades gracefully: when one signal is noisy on a given query, the other usually carries it. The sampling costs you a few extra calls, so reserve the consistency check for high-stakes paths and lean on verbalized confidence for the cheap ones.
Calibrate the threshold against a real error budget
Here is the part teams skip, and it is the part that matters. A confidence score of 0.8 means nothing until you know what error rate it buys you. So you measure it.
Take a few hundred questions where you know the right answer. Run them through your signal. Now you can see, for any threshold, two numbers: coverage, the fraction of questions you actually answered, and selective error, the fraction of the answered ones you got wrong. Those two trade off. Raise the threshold and your answers get more reliable but you abstain more often. Lower it and you answer more but you are wrong more. The business decides which error rate it can live with, and you pick the lowest threshold that stays under it.
def pick_threshold(scored, max_error=0.05):
"""scored: list of (confidence, is_correct) from a labeled dev set.
Return the lowest threshold whose selective error stays under max_error,
which maximizes coverage for the error budget you're willing to accept."""
best = None
# sweep candidate thresholds from the observed scores
for t in sorted({round(s, 2) for s, _ in scored}):
answered = [correct for conf, correct in scored if conf >= t]
if not answered:
continue
error = 1 - (sum(answered) / len(answered))
coverage = len(answered) / len(scored)
if error <= max_error:
# first (lowest) threshold under budget gives the most coverage
if best is None or coverage > best["coverage"]:
best = {"threshold": t, "error": error, "coverage": coverage}
return best or {"threshold": 1.0, "error": 0.0, "coverage": 0.0}
# e.g. {"threshold": 0.82, "error": 0.048, "coverage": 0.71}
# In plain terms: answer 71% of questions, and among those you're wrong 4.8% of the time.
That output is the sentence you can take to a stakeholder. "We answer seventy-one percent of questions and the ones we answer are wrong under five percent of the time" is a real service level. "The model is pretty good" is not. And the questions you did not answer did not become wrong answers. They became a defer, which is a far cheaper outcome than a confident mistake.
Wire it into the agent loop
The gate itself is boring, which is the point. Draft, score, then route.
def answer_with_gate(question: str, context: str, threshold: float) -> dict:
draft = draft_with_confidence(question, context)
# explicit abstention from the model is already a clear signal
if draft.get("answer") is None:
return {"action": "abstain", "message": "I don't have enough to answer that confidently."}
verbal = float(draft.get("confidence", 0.0))
consistency = consistency_score(question, context)
score = 0.5 * verbal + 0.5 * consistency # blend; tune the weights on your data
if score >= threshold:
return {"action": "answer", "answer": draft["answer"], "score": round(score, 2)}
if score >= threshold - 0.15:
# in the gray band, spend more compute rather than guess or give up
return {"action": "escalate", "answer": draft["answer"], "score": round(score, 2)}
return {"action": "abstain",
"message": "I'm not confident enough to answer that. Want me to route you to a human?"}
The three actions matter. Answering is the happy path. Abstaining protects the user from a wrong answer and protects you from the fallout. Escalating is the one people forget, and it is where a lot of the value lives. When the agent is unsure but the question is answerable, you do not have to choose between guessing and giving up. You spend more, either on a bigger model or a human in the loop, on exactly the queries that need it. That is selective prediction and cost control in the same mechanism. The related pattern of deferring to a person on the shakiest cases is worth its own treatment, and I wrote about the workflow side of that in human-in-the-loop approval gates.
Tradeoffs and the things that will bite you
Verbalized confidence is miscalibrated out of the box. A model that says 0.9 is not right ninety percent of the time. That is fine, because you never trust the raw number. You trust the threshold you calibrated against real outcomes, and the raw score only needs to rank answers roughly by reliability, not to be a true probability.
Coverage can collapse. Set the error budget too tight and the agent abstains on everything, which is safe and useless. Watch coverage as closely as error. An agent that answers ten percent of questions perfectly has not solved your problem, it has moved it to whatever catches the other ninety percent.
Thresholds drift across domains and over time. The number you calibrate on billing questions will not hold for technical troubleshooting, and the number that held last quarter will slip when the content or the model changes. Segment your calibration by query type where the volume justifies it, and re-run it on a schedule the way you would any eval. If you are already sampling for the consistency signal, you have most of the machinery for a voting-based reliability check too, which I go deeper on in the piece on consensus sampling and voting.
Sampling has a latency and cost tail. Five calls per answer is real money and real milliseconds. Gate the gate: run verbalized confidence on everything, and only trigger the multi-sample consistency check when the cheap signal lands in the uncertain band. Most queries never pay the full cost.
And do not let your offline evals keep punishing abstention. If your test harness scores an "I don't know" as a failure, you are optimizing against the exact behavior you just built. Give abstention partial credit, or better, score selective error and coverage as separate numbers so a correct decline reads as a win.
The takeaway
Agents guess because we built the whole training and evaluation stack to reward guessing. You will not prompt your way out of that, but you do not have to live with it either. Put an abstention gate in front of the model: a confidence signal you can get from the outside, a threshold you calibrated against an error budget a stakeholder actually agreed to, and a three-way choice to answer, defer, or escalate. The engineering is deliberately unglamorous, and that is why it works. What you get back is an agent that is trustworthy precisely because it knows when to stop talking.
If your agents are answering questions they should be declining, that gap is measurable in a week and fixable soon after, without touching the model. Book a consultation call and we can figure out where your agent should be saying "I don't know" and build the gate that makes it.
Viral Ruparel
Generative AI consultant helping teams ship reliable LLM and agent systems in production.
Contact Viral about your AI project →Frequently Asked Questions
What is an abstention gate for an LLM agent?+
An abstention gate is a small decision layer that sits between the model's draft answer and the user. It reads a confidence signal for the answer, compares it against a threshold you calibrated on labeled data, and then chooses one of three actions: return the answer, abstain and say the agent is not sure, or escalate to a larger model or a human. The point is to stop the agent from returning confident wrong answers by making "I don't know" a first-class outcome instead of something the model was trained to avoid.
Why do LLMs guess instead of saying they don't know?+
Because that is what their training and evaluation rewarded. Most benchmarks score a wrong answer and an abstention the same way, as zero credit, while a lucky guess sometimes scores full marks. Under that scheme the expected value of guessing is always at least as high as abstaining, so models learn to always produce an answer. RLHF can amplify the bias further when human raters prefer long, confident responses over a short "I am not sure." The behavior is rational given the incentives, which is why you fix it at the system level rather than by asking the model to try harder.
How do you measure agent confidence without access to model internals?+
Two signals work well and need no access to logits or weights. The first is verbalized confidence, where you ask the model to rate how sure it is on a fixed scale as part of a structured response. The second is self-consistency disagreement, where you sample the same question a few times and measure how much the answers diverge, since high disagreement is a reliable proxy for uncertainty. Both are model-agnostic, both work through a plain API, and you can combine them. You then calibrate the threshold on a labeled set so the raw score maps to a real error rate.
Related Articles
Nobody Can Tell You How Many Agents You Are Running
A year of shipping agents leaves most teams with a fleet nobody can inventory: no owner, no declared budget, no record of which tools each one can reach. That is agent sprawl, and it shows up as a surprise invoice and an incident with no name on it. The fix is boring and it works: an agent registry. A typed manifest every agent must declare, a CI gate that blocks anything unregistered or over-budget, and a reconciliation job that flags an agent doing more than it said it would.
Your Agent's Memory Is Full and Most of It Is Junk
You gave your agent persistent memory and it worked. Six months later the store is full of duplicates, contradictions, and vague paraphrases that crowd out the facts you actually need, and retrieval quietly gets worse every week. This is a garbage-collection problem, not a storage problem. Here is how to build the curation layer: gate what gets written, deduplicate and reconcile on the way in, decay what stops earning its slot, and measure whether the store is still healthy.
Your Agent Fetches the Same Row Ten Times a Session
A long-running agent calls the same read tool over and over inside a single session, paying full latency and quota for answers it already had a few turns ago. A tool result cache fixes it, but the naive version ships stale data and quiet correctness bugs. Here is how to build one that classifies which tools are cacheable, deduplicates concurrent calls with singleflight, and invalidates reads the moment a write touches the same data.