Prompt caching explained: why cached tokens cut your AI bill
Coding agents re-send the same context on every turn. Prompt caching bills those repeated tokens at a fraction of the input price — here's how it works and how to benefit.
If you look at the requests a coding agent makes, you'll notice something: most of every prompt is the same as the previous one. The system prompt, tool definitions, the files already read, the conversation so far — all re-sent, turn after turn.
Prompt caching exists so you don't pay full price for that repetition.
How it works
When a model provider sees a prompt prefix it processed recently, it can reuse the internal state instead of recomputing it. Those tokens are reported separately:
- OpenAI-style APIs report them as
cached_tokens, - Anthropic-style APIs as
cache_read_input_tokens(andcache_creation_input_tokenswhen a new cache entry is written).
Cached tokens are billed at the cache price, typically around 10% of the normal input price.
A quick example
Say an agent session makes 200 requests with an average prompt of 40,000 tokens, of which 34,000 are an unchanged prefix:
| Tokens | Without cache | With cache | |
|---|---|---|---|
| Unchanged prefix | 6.8M | billed as input | billed at cache price |
| New input | 1.2M | input price | input price |
With a cache price at one tenth of input, the prefix costs 90% less — which usually makes caching the single biggest factor in what an agent session costs.
Getting the most out of it
- Keep the beginning of your prompt stable. Put the system prompt and tool definitions first and don't change them between calls.
- Append, don't rewrite. Agents that add to the conversation benefit more than ones that re-summarise it every turn.
- Stay on one model per session. Caches are per model — switching throws the prefix away.
- Check your numbers. The usage page shows cached tokens and the savings for every period.
Is caching automatic?
For most models, yes: the provider decides what to cache and you simply see cheaper cached tokens in the usage. Anthropic-format requests can also mark cache breakpoints explicitly with cache_control — Claude Code does this for you.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key