Skip to content
Apixo
Blog
blog· 2 min read

Prompt caching explained: why cached tokens cut your AI bill

Coding agents re-send the same context on every turn. Prompt caching bills those repeated tokens at a fraction of the input price — here's how it works and how to benefit.

If you look at the requests a coding agent makes, you'll notice something: most of every prompt is the same as the previous one. The system prompt, tool definitions, the files already read, the conversation so far — all re-sent, turn after turn.

Prompt caching exists so you don't pay full price for that repetition.

How it works

When a model provider sees a prompt prefix it processed recently, it can reuse the internal state instead of recomputing it. Those tokens are reported separately:

  • OpenAI-style APIs report them as cached_tokens,
  • Anthropic-style APIs as cache_read_input_tokens (and cache_creation_input_tokens when a new cache entry is written).

Cached tokens are billed at the cache price, typically around 10% of the normal input price.

A quick example

Say an agent session makes 200 requests with an average prompt of 40,000 tokens, of which 34,000 are an unchanged prefix:

Tokens Without cache With cache
Unchanged prefix 6.8M billed as input billed at cache price
New input 1.2M input price input price

With a cache price at one tenth of input, the prefix costs 90% less — which usually makes caching the single biggest factor in what an agent session costs.

Getting the most out of it

  1. Keep the beginning of your prompt stable. Put the system prompt and tool definitions first and don't change them between calls.
  2. Append, don't rewrite. Agents that add to the conversation benefit more than ones that re-summarise it every turn.
  3. Stay on one model per session. Caches are per model — switching throws the prefix away.
  4. Check your numbers. The usage page shows cached tokens and the savings for every period.

Is caching automatic?

For most models, yes: the provider decides what to cache and you simply see cheaper cached tokens in the usage. Anthropic-format requests can also mark cache breakpoints explicitly with cache_control — Claude Code does this for you.

#caching#pricing
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key