•11 min
You Pay Full Price for the Same 8,000-Token Prompt on Every Agent Step
An agent resends the same enormous system prompt, tool list, and history on every step, and without prefix caching you pay full input price and full prefill latency each time. Structuring the prompt so the provider reuses its KV cache cuts both, and one wrong byte quietly turns it all off.
LLM
Agents