•10 min
Your Model Advertises 1M Tokens. It Starts Forgetting Around 600K.
A million-token context window landed in half the models this week, and the reflex is to stop retrieving and just paste everything in. The window is real. The quality across all of it is not. Here is how to measure your model's actual effective context, then spend the window with a token budget instead of filling it and paying linearly for output that quietly gets worse.
LLM Engineering
Context Engineering