Your Model Router Saves Pennies and Burns Your Prompt Cache

You added a model router to cut costs and the bill went up. The savings were never in the sticker price, they were in the prompt cache, and a router that re-picks the model every turn throws that cache away on each switch.

Viral RuparelNo. 64 of 6411 min read

A team I was helping had done everything the cost playbook told them to. They had put a router in front of their agent that sent easy turns to a small model and hard turns to a flagship, and on paper the blended token price dropped by half. A month in, the bill was higher than before the router existed. Not by a little. Every invoice line looked cheap, the per-call prices were exactly what they expected, and the total still climbed.

The money was never in the sticker price. It was in the cache. Their agent, like most agents, resends a large and nearly identical prompt on every step: the system instructions, the tool schemas, the running history. When that prefix stays on one model the provider serves it from its own prefill cache at a small fraction of the input price. The router they added re-picked the model on every turn, and a cache that belongs to one model is useless to another. Each switch paid a fresh cold write where it could have paid a cheap read. They had carefully optimized the one number on the invoice that was already small and quietly wrecked the one that was large.

Per-turn routing writes a cold cache on every model switch; session routing keeps one warm cache with a single gated escalation.
Figure. Routing against the cache, and with it

The savings that were never where you thought

This is the collision nobody warns you about. The cost-cutting move everyone reaches for first is routing calls to the cheapest capable model, and the second is structuring the prompt so the stable prefix caches. Done separately they both work. Wired together without care, the first one eats the second.

The reason it hides is that the two systems report success in different places. Routing proves itself on the per-call price, which is exactly what the invoice shows you line by line, so the router looks like it is working. Caching proves itself in a single ratio buried in the usage object, the share of prefix tokens served as reads, which nobody is looking at. So the visible number improves while the invisible one collapses, and the net comes out worse. The rest of this essay is about seeing the invisible number and arranging your routing so it does not wreck it.

Why the cache is the real line item

Start with the arithmetic, because it is lopsided in a way that is easy to miss. A cache read on current models costs roughly a tenth of the base input price, and as little as a twentieth on the larger models, while a cache write costs more than a full-price request: about 1.25 times for the short-lived entry and about 2 times for the hour-long one, per the Claude prompt caching documentation. So a warm read is on the order of ten to twenty times cheaper than paying for those tokens fresh, and the gap between a read and a write is wider still.

Now picture the shape of an agent. It sends a twenty-thousand-token prefix on step one, then adds a few hundred new tokens per step over thirty steps. The new tokens are a rounding error. The prefix, resent thirty times, is the bill. Whether each of those resends lands as a read or a write is therefore the single largest lever you have, larger than which model answered. A recent evaluation across OpenAI, Anthropic and Google, "Don't Break the Cache", measured 41 to 80 percent cost reductions and 13 to 31 percent faster time to first token from caching on long-horizon agentic tasks, and found that how you arrange the prompt matters more than whether caching is simply turned on.

The detail that makes routing dangerous is one line in that documentation: caches are scoped to the model. The saved prefill state belongs to the model that produced it, and no other model can read it. Switch models and you do not get a smaller discount, you get none. The cache you were relying on for most of your savings is gone the instant the router points somewhere else.

What a per-turn router actually does to the cache

Here is the pattern that backfires. The router runs on every turn, classifies the incoming request, and picks a model for it in isolation.

# Per-turn routing: looks cheap on the invoice, runs cold underneath.
def answer(turn, history):
    model = classify(turn)              # a fresh decision every single turn
    messages = history + [turn]
    resp = call(
        model=model,
        system=SYSTEM_PROMPT,           # large, stable, meant to be cached
        tools=TOOLS,                    # also large and stable
        messages=messages,             # append-only, grows each step
    )
    history.append(resp)
    return resp
# The prefix (system + tools + history) is identical turn to turn,
# but it only caches against the model that last saw it. Any time
# `classify` returns a different model, that prefix is re-written cold.

A real conversation rarely stays in one difficulty tier. It drifts: an easy question, a hard follow-up, two more easy ones. The router faithfully sends those to small, flagship, small, small, and the prefix is written cold on the flagship, then written cold again when it returns to the small model, because by then the small model's entry may have been evicted or was never large enough to matter. You are not caching. You are paying the write premium on nearly every step and collecting almost no reads. The invoice shows a cheap model doing cheap work, which is exactly why this hides for a month.

Providers do not share a cache, and they do not agree on the rules

It gets worse the moment routing crosses a provider boundary, which it does the instant you add failover across providers for resilience. Two things break at once.

First, the caches are in separate universes. A prefix warm on one vendor's model does not exist on another's. A failover is a guaranteed cold start for every live session at once.

Second, the providers do not even agree on how caching works, so you cannot reuse your prompt layout across them. Anthropic uses explicit cache_control breakpoints, up to four per request, with a five-minute default lifetime or an hour at higher write cost, and reads that look back over the recent blocks. OpenAI, by its prompt caching guide, caches automatically with no code changes, matches on the exact rendered prefix, and retains entries for about thirty minutes on its newer models. The prompt you tuned so the breakpoint sits on the last stable block, the discipline that the prefix-caching work is all about, has no equivalent on the provider that just caches whatever prefix repeats. Your carefully placed breakpoints are inert there, and the exact-match provider will miss on a prefix that your other provider happily reused. Failover is not a drop-in. It is a cold write plus a silent reformat, and if you measured cache hits only on your primary you will not see the second provider running cold until the bill arrives.

Route the session, not the turn

The fix is to move the routing decision up a level, from the turn to the session. Decide the model once, keep the whole conversation on it, and treat a model switch as a deliberate exception rather than a per-turn reflex. That is cache affinity: the thing that keeps a prefix warm is that it keeps hitting the same model.

# Session-sticky routing: decide once, then let the cache do its job.
def route(session):
    if session.model is None:
        session.model = classify(session.first_turn)   # the one up-front call
    return session.model

def answer(session, turn):
    model = route(session)
    resp = call(model, SYSTEM_PROMPT, TOOLS, session.history + [turn])
    session.history.append(resp)
    if needs_escalation(resp):          # the single controlled switch
        session.model = FLAGSHIP        # pay one cold write, then stay put
    return resp
# Most sessions make exactly one routing decision and ride a warm cache
# for the rest of their life. The expensive switch happens only when a
# real signal, not a per-turn classifier, says the small model is stuck.

Notice what this gives up and what it keeps. You lose the ability to shave a few cents on an easy turn in the middle of a hard session. You keep the cache warm for every turn in between, which on an agent is worth far more. The predictive router still earns its place at the top of the session, where it commits you to a model before the cache exists, and the escalation gate still protects the hard tail. You have just stopped letting the router churn the one asset that was paying for everything.

Make the cache a term in the routing math

Once routing is a session decision, the remaining switches, the escalations and the failovers, are the ones worth doing arithmetic on. The question is never only "is this model cheaper per token." It is "is this model cheap enough to also cover the cold write I have to pay to move to it, given how many more turns this session will run."

# Should we actually switch models, cache cost included?
def worth_switching(prefix_tokens, remaining_turns, current, candidate):
    # Staying: warm reads on the model that already holds the prefix.
    stay = prefix_tokens * current.read_price * remaining_turns
    # Switching: one cold write now, then warm reads on the new model.
    switch = (prefix_tokens * candidate.write_price
              + prefix_tokens * candidate.read_price * (remaining_turns - 1))
    return switch < stay
# A cheaper model late in a long session often still loses, because the
# write wipes out the per-token saving when few reads remain to amortize it.

You cannot run this calculation on guesses, which means the whole approach rests on measurement. Log the cache-read and cache-write token counts on every response, tagged with the model and provider that served it, and watch the share of prefix tokens served as reads. That number is your real cost health, and it belongs on a dashboard next to spend. Wire it through your OpenTelemetry traces so the cache read ratio is visible per route, and tie the spend side into your per-tenant cost budgets so a tenant whose sessions keep churning the cache shows up as a cost anomaly rather than a surprise. For the subset of traffic that repeats near-identical questions across sessions, a semantic cache in front of the model sidesteps the whole routing question by not calling a model at all.

The parts that bite

Sticky, cache-aware routing removes the common own-goal, but it introduces edges of its own that are worth naming plainly.

Failover storms are the sharpest one. A provider blips, the circuit breaker trips, and every live session fails over to the backup at the same moment. Each of those sessions takes a cold write on the backup provider simultaneously, so the mechanism meant to protect you from an outage hands you a synchronized cost spike on top of it. Budget for that write, and stagger re-entry when the primary recovers rather than stampeding back.

The hour-long cache is a two-times write, so pre-warming is not free insurance. Pre-warm the wrong prompt, or pick the long time-to-live for traffic that actually arrives every couple of minutes, and you pay the premium to prime a cache you would have filled cheaply anyway.

Sticky routing can strand the hard tail. A session that starts easy and turns hard will sit on the small model until something forces a switch, so the escalation gate is not optional. Make it fire on a real signal, a failed check or a low-confidence answer, and keep it rare. This is the same discipline as matching reasoning effort to the difficulty of the request: spend the expensive resource only where it changes the outcome.

Cross-provider measurement is uneven. Anthropic reports cache_read_input_tokens and cache_creation_input_tokens, OpenAI reports cached_tokens, and others differ again. Normalize them into one schema at ingestion or your dashboard will quietly compare unlike numbers and tell you the cache is fine when it is cold on half your traffic.

Zero Data Retention changes the whole calculation. If your organization runs under ZDR, some providers disable prompt caching outright, and a cost model built on warm reads no longer holds. Confirm caching is actually available on your account before you design a routing strategy around it.

The takeaway

The instinct to cut LLM cost by routing to cheaper models is sound, and so is the instinct to cache the giant repeated prefix. The trap is that they share a boundary, the model, and a router that moves the boundary every turn destroys the cache that was doing most of the saving. The sticker price is the small number. The cache is the large one. Optimize them together or you will proudly drive the first one down while the second one swallows the win, and you will not see it until a month of invoices has stacked up.

So route the session, not the turn. Keep each conversation on one model, keep its prefix warm, make the few real switches earn the cold write they cost, and measure the read ratio per provider so you are never guessing. If your bill climbed right after you added a router or a failover path and the per-call numbers all look cheap, the cache is almost certainly where it is leaking, and it is the kind of thing that is quick to find once you know to look. If that sounds like your setup, book a short call and we can trace where your prompts are running cold.

ShareXLinkedInviralruparel.com/blog/prompt-cache-model-routing-provider-failover

Questions this essay answers

01Does model routing reduce or increase LLM costs?

It can do either. Routing lowers the per-call sticker price by sending easy work to cheaper models, but for an agent that resends a large prefix every step, most of the real bill is prompt-cache reads, not fresh tokens. A router that re-picks the model each turn fragments that cache across models, so every switch pays a cold write instead of a cheap read. If the cache savings you lose are larger than the sticker savings you gain, routing makes the total go up.

02Why does switching providers reset the prompt cache?

Prompt caches are scoped to a single model inside a single provider. Each provider keeps its own cache namespace and its own matching rules, so a prefix that is warm on one model is simply not present anywhere else. A failover to a second provider is therefore a guaranteed cold write, and often a reformat too, because the prompt structure you tuned for one provider's breakpoints does not line up with another provider's exact-prefix matching.

03What is cache-aware or sticky routing?

Sticky routing decides the model once, usually at the start of a session, and keeps the whole conversation on it so the cached prefix is reused on every following turn. Cache-aware routing goes one step further and treats the cost of a cold cache write as part of the routing decision, only switching models when the expected savings clearly beat the write you have to pay to get there.

04How do I measure whether routing is hurting my cache?

Log cache-read and cache-write token counts on every response, split by the model and provider that served it. The number that matters is the share of prefix tokens served as reads rather than writes. If that ratio fell after you added a router, the router is churning the cache. Providers name these fields differently, so normalize them into one schema before you put them on a dashboard.

Working on this in production?If your bill climbed after you added a router or a failover path, the cache is usually the leak. Thirty minutes, no pitch, and we find where it is cold.
Viral Ruparel

Generative AI consultant and full-stack architect. Ex-IBM, former Director of Engineering at a YC W20 startup. Writes about what breaks when agents meet reality.