LLM Cost Optimization for Agents: A Production Guide
Every lever that cuts an agent's LLM bill, in the order they pay off: routing, effort, prefix and result caching, context control, per-tenant caps and tracing.
Every agent bill I have looked at had the same shape. A handful of levers nobody had pulled, a few that had been pulled in a way that quietly did nothing, and no single number that said where the money went. The invoice moved with traffic, with how chatty the model felt that week, with one customer's data, and the team was reading it after the fact with no way to tie a dollar back to a decision.
That is because an agent's cost is not a flat line the way a REST endpoint's is. It is elastic and path-dependent. The same task runs for two dollars one day and forty the next, because on the second day a tool failed, the agent retried, re-read the same files, and looped eleven times before it finished. Every turn of that loop resent the whole conversation: the system prompt, the tool definitions, every prior tool result. Input tokens are where agent spend goes, and an agent is a machine that resends almost the same enormous prompt on every step.
This guide is the calm reference next to the essays it draws on. It walks the levers in the order they tend to pay off: match the model and the thinking to the request, make the calls you still make cheaper, avoid the calls you do not need, keep the context from compounding, stop paying for dead time and abandoned work, cap what any one tenant can spend, and trace where it all went.
Match the model and the thinking to the request
Real traffic is lopsided. Support bots spend most of their day on FAQ-shaped questions, coding assistants on small edits, extraction pipelines on work that is easy by definition. The hard tail is real, but it is the minority. If every request goes to your most expensive model, every password reset pays a flagship premium for output a model costing a fifth as much would have nailed.
Routing fixes that without touching the quality of the hard requests, because those still go to the same model they always did. A predictive router uses a small, cheap model to classify the request into a tier, simple, moderate or hard, and a registry maps each tier to a concrete model in one place. A cascade tries the cheapest model first, checks the answer with a separate grader or, better, a deterministic check like code that compiles or JSON that parses, and escalates only on failure. Use the router when difficulty is visible in the request and latency matters. Use the cascade when correctness is checkable after the fact. Log the chosen model and the outcome so the thresholds can move. Bias those thresholds so the failure mode is an unnecessary escalation, not a confident wrong answer from a weak model. I went through the registry, the classifier, the cascade and the traps in model routing to cut AI costs.
The same logic applies inside a single model. On a reasoning model, thinking tokens are billed at the output rate and are usually the largest and most variable slice of the bill. Turn extended thinking on globally and you have bought every easy request a slow reasoning pass it never uses. Rough arithmetic on 200,000 requests a day where a fifth are genuinely hard puts about three quarters of the thinking spend on requests that did not need it. The fix has the same shape as routing: a cheap gate, a heuristic first and a small non-reasoning model only for the ambiguous middle, picks an effort tier per request, and the tier maps onto whatever knob the provider exposes. On current Claude models that is an effort setting from low through max on adaptive thinking; the old fixed budget field is rejected. Put a soft daily ceiling on top that demotes borderline requests when the budget runs low and never demotes a genuinely hard one. Measure accuracy per difficulty tier, not as one blended number, because an average hides a regression on the hard twenty percent. Reach for effort before you build a multi-model cascade, since it keeps one cache namespace and one behavior profile. The classifier and the ceiling are in reasoning effort budgeting.
One rule covers both: never let the classifier cost more than the work it is steering.
Cache the prefix the provider already processed
Look at what changes between two consecutive agent steps. The system prompt is byte-for-byte identical. The tool schemas are identical. The first twenty-nine messages are identical. The new content is one tool result and the model's next move. You are resending fifteen thousand stable tokens to append two hundred new ones, and without prompt caching you pay full input price and full prefill latency for all of them, on every step.
Prompt caching is a prefix match on the exact bytes of your request up to a breakpoint. The provider saves the attention state for that prefix and reuses it when the next request starts with the same bytes. Cache reads bill at roughly a tenth of the base input price, and the prefill for that span mostly disappears. It never changes the output, which makes it the one cost lever with no quality risk and the first one to reach for.
The one rule everything follows from: any change anywhere in the prefix invalidates the cache for everything after it. So the stable content goes first and the volatile content goes last. The classic mistake is a timestamp interpolated into the system prompt, which produces different leading bytes on every request and turns caching off.
# Stable prefix first, marked cacheable. Volatile data goes into messages,
# after the breakpoint, where it invalidates nothing ahead of it.
response = client.messages.create(
model="claude-opus-5",
max_tokens=2048,
system=[{
"type": "text",
"text": RULES, # frozen: no timestamps, no user id, no request id
"cache_control": {"type": "ephemeral"},
}],
tools=TOOLS, # deterministic order, same every step
messages=messages, # append-only; the newest turn carries a breakpoint
)
u = response.usage
total = u.cache_read_input_tokens + u.cache_creation_input_tokens + u.input_tokens
print(f"cache hit rate: {u.cache_read_input_tokens / total:.0%}" if total else "n/a")
For an agent there is a second discipline: treat the sent history as append-only. Roll the breakpoint forward onto the newest turn each step, and never reach back to edit an earlier message, because the moment you rewrite history you pay full price to reprocess everything from the edit forward. To inject a mid-run instruction, append a system-role message to the end rather than editing the top-level system prompt. Then read the usage numbers on every response, because adding a cache marker is not the same as caching. Writes cost about 1.25x a normal request, the default entry lives about five minutes after its last use, prefixes below a model-dependent minimum do not cache at all, and the cache is scoped to the model, so switch models on your main loop and you lose it. All of that, with the list of silent invalidators to grep for, is in prompt prefix caching for agents.
Avoid the calls you have already made
Prefix caching makes the calls you still make cheaper. The next two caches remove calls entirely, at two different layers.
Pull a month of agent logs and group requests by meaning rather than by string. The same handful of questions comes up all day in different words, and your agent reasons about each one from scratch. A semantic cache stores each answered question as an embedding next to its response, and on a new question it finds the nearest stored vector and returns the stored answer if the similarity clears a threshold. The threshold is the whole game: too strict and you have exact matching with extra steps, too loose and you serve the answer for "cancel my subscription" to someone who asked to pause it. Start around 0.9 on a strong embedding model, log every near-miss with its score, and move the line only with labeled pairs in front of you. Scope the key so that everything that changes the correct answer, tenant, locale, tool version, is part of the entry's identity, or you will serve one user's account details to another. Never cache anything time-sensitive, stateful or side-effecting, give every entry a TTL sized to how fast the truth moves, and wire targeted invalidation to the events that make a class of answers stale. Measure your repetition rate before you build this. The guardrails are in semantic caching for LLM agents.
The other cache sits in your own runtime, between the agent and its tools. Attach a counter to the tool layer and you will find the same user record pulled on turn 2, turn 5 and turn 9. Each one is a full round trip for an answer that has not changed. Unlike the semantic cache, this one keys on the tool name plus normalized arguments, so a hit is an exact repeat, never a guess. Cacheability is a property you declare per tool and it defaults to off; only genuine reads opt in, each with a TTL that bounds how stale an answer you are willing to act on. Every cached read is tagged with the entity families it depends on, and every write declares the families it mutates, so a write clears exactly the reads it affects before the next read happens. Add singleflight so ten concurrent requests for the same uncached key become one request, which matters most when an agent fans out tool calls in parallel. Invalidation only covers writes that go through your agent, so keep TTLs short for anything with external writers. The policy, the key, the singleflight and the executor are in tool result caching with singleflight and invalidation.
Keep tool definitions and tool output out of the window
Two things bloat an agent's context in ways that have nothing to do with the work it does.
The first is the tool catalog. Connect a dozen MCP servers and the model holds every tool's name, description and full JSON schema on every turn, before the user has typed a character. On a tool-heavy setup that is well over a hundred thousand tokens of overhead, scaling with how many integrations you connect rather than how much work gets done. The second is worse because it scales with your data: when an agent chains two tools, every intermediate result passes through the model twice. A two-hour call transcript moving from Drive to a CRM is fifty thousand tokens the model does not need to read, only to carry.
The code execution pattern fixes both. Instead of a menu of forty definitions, you give the model a sandbox and a thin catalog of importable modules, one per server, and it writes a short script that calls exactly the tools it needs. Schemas resolve at import time, so a fiftieth server costs one line in the catalog, not another schema in every prompt. Bulk data stays in the sandbox and only the distilled result crosses back.
// The model writes and runs this inside the sandbox.
import { query } from "./servers/salesforce/query";
const rows = await query({
soql: "SELECT Amount FROM Opportunity WHERE StageName = 'Closed Won'",
});
// 4,000 rows stay in the sandbox. One sentence crosses back to the model.
const total = rows.reduce((sum, row) => sum + row.Amount, 0);
console.log(`Closed-won pipeline: ${rows.length} deals, $${total.toLocaleString()}`);
Anthropic's own writeup reports a workflow dropping from about 150,000 tokens to 2,000 this way. The price is that you are now executing model-generated code, so you need a real sandbox with no ambient credentials, tight resource limits and a locked-down network. It is overkill for an agent with three tools and small payloads. The tradeoffs are in MCP code execution and agent token costs.
If you are not ready for a sandbox, the smaller version of the same idea works on any tool loop. A single tool call can return 38,000 tokens of account JSON when the task needed one boolean in field 14, and because the blob sits in the message history, every later turn pays for it again. Offloading intercepts large results at the tool boundary: a wrapper measures the output, passes small results through inline, and writes large ones to a store, returning a handle plus a preview. The preview is the part that matters. For structured data it is the shape; for text it is the first and last lines and a total count. A companion read tool takes the handle and a key path or a line range, so the model fetches only the slice it needs. A thin preview is worse than no offloading, do not offload what the model needs on every turn, and a handle to a resource a later write changed is a stale snapshot, so key artifacts by the resource they describe. The wrapper, the preview builder and the read tool are in tool output offloading.
Compact the history before it compounds
Offloading stops the blobs entering the window. The running conversation still grows, and every turn resends all of it. A conversation that reaches 150,000 tokens does not cost 150,000 tokens once; it costs that on this turn, and the next, and the one after. At Opus pricing an agent averaging 120,000 tokens of context across fifty turns spends around thirty dollars on input before it writes a single output token. Time to first token tracks prompt size too, and a long messy history causes drift: the model loses track of what it decided, repeats finished work, contradicts itself.
The lazy fix is truncation, which is the fastest way to make the agent forget the plan it committed to on turn three. The pattern that holds up is anchored summarization: one persistent summary of everything old, the last four to eight turns kept verbatim, and when the total crosses a threshold, fold the turns about to age out into the anchor rather than rebuilding the summary from scratch. Use the provider's token-counting endpoint to decide when, not a client-side guess. Spell out in the summarizer's prompt what must always survive: decisions, IDs, file paths, open TODOs. Compact only at turn boundaries, never between a tool_use block and its tool_result. And put the anchor where it caches well, appended to the end of the stable system prompt, because splicing it in at the front invalidates the prefix cache every time it changes and quietly undoes a chunk of the savings you were chasing.
If you would rather not maintain summarization logic, the Anthropic API has server-side compaction in beta behind a header and a context-management directive. The one rule that trips everyone: append the full response content back to the message list, compaction blocks included, or the feature silently does nothing. Both versions are in context compaction for long-running agents.
Stop paying for dead time and dead work
Two runtime habits waste money without a single wasted token in the prompt.
The first is serial tool execution. When the model can see that several tool calls are independent, it puts all of them in one assistant turn. That grouping is the model telling you these calls can run at the same time. Most loops throw it away by awaiting each call in turn, so three lookups of 400, 500 and 300 milliseconds take 1,200 milliseconds instead of 500. Run the batch concurrently with asyncio.gather and return_exceptions=True, so one failing call does not cancel its siblings, then hand back a result for every tool_call_id, with a structured error in the slot that failed, so the model can retry just that one. Cap the fan-out with a semaphore, because a model that emits forty independent calls does not know your rate limits. Keep genuine writes serial, and never parallelize across turns on your own, since the independence guarantee only holds inside the batch the model gave you. The loop and its pitfalls are in parallel tool execution.
The second habit is finishing work nobody is waiting for. A user closes the tab after eight seconds; on the server nothing notices. The agent finishes its model call, fires three tool calls, spawns a subagent, and completes forty seconds later with an answer for nobody. A single outer timeout does not help, because it fires on the top coroutine while every leaf below it keeps running. You stopped listening, not spending. The fix is an absolute deadline, a single point on the monotonic clock computed once at the top and passed unchanged all the way down, plus a cancel that propagates through every tool call and subagent and tears in-flight siblings down the moment one fails or the clock runs out. Cleanup gets its own small, shielded time budget, because the deadline that triggered it has zero time left. Wire the client disconnect to cancelling the top-level task and the same teardown path handles both cases. Irreversible side effects belong behind a checkpoint or an approval gate, not behind cancellation. I went through the deadline object, the propagation and the teardown in deadline and cancellation propagation.
Cap spend per tenant before the call
Everything above lowers the average. None of it caps the total. A cheaper model does not stop a tenant making a hundred thousand calls. A rate limiter counts requests, not dollars, and ten cheap requests look identical to ten that each drag fifty thousand tokens of context. One customer who points your agent at their entire document store on a Tuesday can spend the month's budget by 2pm with nothing broken, because nothing in the request path was counting dollars and nothing was allowed to say no.
The fix is a ledger and a gate. Cost is token counts times a per-model price, so keep the prices in one table and compute each call's cost from the usage the response reports. The ledger increments a tenant's spend for the current window atomically, a Redis INCRBYFLOAT on a key that expires with the window, because a read-modify-write in application code races exactly when the cap most needs to hold. The detail that separates a guardrail that holds from one that leaks is the order of operations. Before the call you do not know the output length, and output tokens are the expensive ones, so reserve the worst case, counted input plus max_tokens priced as output, check it against the budget, and commit the reservation before the call goes out. After the call, settle to the real figure.
def guarded_call(tenant_id, model, messages, max_tokens, daily_budget):
window = date.today().isoformat()
# 1. Reserve the worst case: input is known, output is capped at max_tokens.
reservation = cost_usd(model, Usage(count_input_tokens(model, messages), max_tokens))
if current_spend(tenant_id, window) + reservation > daily_budget:
raise BudgetExceeded(tenant_id)
# 2. Commit before the call, so concurrent callers see the money as spent.
record_spend(tenant_id, window, reservation, DAY_TTL)
try:
resp = call_model(model=model, messages=messages, max_tokens=max_tokens)
except Exception:
record_spend(tenant_id, window, -reservation, DAY_TTL) # never billed
raise
# 3. Settle: refund the gap between the reservation and the real cost.
actual = cost_usd(model, Usage(resp.usage.input_tokens, resp.usage.output_tokens))
record_spend(tenant_id, window, actual - reservation, DAY_TTL)
return resp
When the gate trips, blocking is the crude option. For a paying customer mid-conversation, degrade first: walk a ladder from the flagship down to the cheapest model, shrink context, drop optional tool calls, and hard fail only when even the cheap path does not fit. Make the degraded behavior explicit and tiered, not a silent quality drop. Meter every model call inside the run against the same tenant key, not just the top-level request, or the loop's fan-out undercounts you. Decide fail-open versus fail-closed per tenant tier on purpose, and meter a monthly window alongside the daily one if you sell a monthly plan. The ledger, the reserve-then-settle flow and the degrade ladder are in per-tenant LLM cost guardrails.
See where the money went
You cannot tune any of this blind, and plain logs will not show you. Two engineers run the same agent on the same task; one spends two dollars, the other forty, because the second run hit a failing tool, retried, re-read the same files and looped eleven times. What the agent produced is a tree, and you cannot debug a tree one line at a time.
Distributed tracing fits the shape. The whole run is a root span, each loop iteration is a step span, and each model call and tool call nests inside its step. OpenTelemetry's GenAI semantic conventions give you an agreed vocabulary: gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons. Instrument the single model call first, in one helper, so the attribute names live in one place. Nest steps and tools so one run is one tree. Then convert tokens to dollars in a span processor using the same one-place price table, so cost rolls up to the run and the forty-dollar run sorts to the top. Record cached input tokens as their own attribute and price them separately, or a cache-hit-rate regression will look like a cost regression and you will chase the wrong thing.
Keep prompt and completion text opt-in, behind a debug flag or on a sampled fraction, because it is user data and the largest thing in the span. Use tail sampling, which decides after the run finishes, so you keep every run that erred, ran long or crossed a cost threshold; head sampling throws the expensive outliers away at random. Keep high-cardinality values on spans, not on metric dimensions. The wiring is in agent observability with OpenTelemetry GenAI tracing. Tracing tells you where the money went; the per-tenant cap stops it going there next time.
Where to start
If you have none of this, do it in this order. Each step makes the next one measurable.
- Instrument first. Wrap the model call with the GenAI attributes, nest steps and tools into one trace per run, and price the spans. A day of traces tells you whether your bill is prefix resends, tool output, a reasoning setting, a looping step or one tenant.
- Fix the prefix. Freeze the system prompt and tool list, push everything volatile after the breakpoint, treat history as append-only, and confirm cache_read_input_tokens is carrying most of your prefix. A few lines of placement, no quality risk.
- Put a cap in the request path. An atomic per-tenant ledger with reserve-then-settle and a degrade ladder. Until this exists, every improvement below is an average, and the average is not what hurts you.
- Match effort and model to the request. Measure the strong model at low effort before you build a cascade. Then add a router or a cascade for the easy majority, logged so the thresholds can move.
- Keep the window small. Offload large tool outputs behind a handle and a real preview. If the tool catalog itself is the problem, move to code execution with a proper sandbox. Add anchored compaction once runs get long enough to need it.
- Remove the repeats. A tool result cache with declared cacheability, TTLs, singleflight and write invalidation. Then, if your logs show real repetition, a semantic cache with a conservative threshold and scoped keys.
- Fix the runtime habits. Run the model's batch of tool calls concurrently under a semaphore, and pass one absolute deadline and a propagating cancel through every tool and subagent so abandoned runs stop spending.
None of these needs a different model or a rewrite of your agent. They are wrappers at boundaries you already own: the prompt builder, the tool executor, the request handler, the ledger in front of the call. Done in that order, each one shows up as a measured drop in the traces from step one, and the bill stops being something you explain after the fact and becomes something you decided.