Because the model is stateless, every request re-sends the whole conversation, and the front of it, the system prompt and tool definitions, barely changes from call to call. Reprocessing that identical prefix each time is wasteful, so providers offer a prefix cache: they remember the work done on a prefix they have seen before and skip straight to the new part.
How it actually helps
The cache keys on the exact leading tokens of your request. If the first stretch of tokens matches a previous request, the provider reuses the computed state for them:
- It is cheaper. The cached prefix is billed as cache tokens at a reduced rate instead of full input tokens.
- It is faster. Skipping the reprocessing of a large stable prefix cuts the time to first token, which you feel most on long sessions.
The catch: order and stability
The cache only helps for an exact prefix match, and it matches from the very start. That has a real design consequence: put the stable stuff first and the volatile stuff last. If you change something near the top of the prompt, or shuffle the order of what you send, you invalidate the cache from that point on and pay full price again.
Turning it on
Mark the stable front of your prompt as cacheable. On later requests the provider serves it from cache instead of reprocessing it.
const response = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 1024,
system: [
{
type: 'text',
text: longStableSystemPrompt,
cache_control: { type: 'ephemeral' },
},
],
messages,
})Related terms
Cache tokens
Cache tokens are input tokens served from the prefix cache at a reduced rate. They are how prompt caching shows up as a separate line in your usage numbers.
Read definition →ConceptInput tokens
Input tokens are the tokens you send in a request: the system prompt, the conversation history, loaded files, and tool definitions. You are billed for them, and they count against the context window.
Read definition →ConceptStateless
Stateless means the model API keeps no memory between requests. Each call starts blank, so every request must carry all the context the model needs. This is foundational to how agents are built.
Read definition →Explore it visually
- Attention variants: MHA, GQA, MQA, MLA and sliding windowsPress play and watch a KV-cache warehouse fill an H100: full attention needs 68.7 GB at Llama 3.1 8B’s full context, grouped-query attention cuts it to 17.2 GB, one shared head to a thirty-second, DeepSeek-V3’s latent attention compresses 71 times, and a sliding window caps it at 537 MB, all computed from the models’ published configs.
- Prompt Caching VisualisedChange a prompt and trace exactly which cached blocks survive. Explore shared prefixes, block size, expiry and cache isolation in an interactive 2D and 3D prefix tree.
- Prompt caching: reuse the stable prefixA support bot resends the same 5,897-token prefix on every call. Watch real Claude Sonnet 5 calls read it back from a prompt cache for a tenth of the price, see a question-first prompt reuse nothing, and see one timestamp at the top of the system prompt throw the whole cache away.