When a prefix cache hits, the tokens it covers do not vanish from your bill; they get reclassified. Cache tokens are the portion of your input tokens that the provider served from cache rather than processing fresh. Same tokens, cheaper rate.
Reading them in your usage
This is mostly something you notice in the numbers. Providers break input down into a few buckets in the usage they return on each request, and you will typically see something like:
- Cache read tokens. Prefix tokens that were already cached and reused, billed at a large discount.
- Cache write tokens. Tokens being stored into the cache for the first time, sometimes billed at a slight premium.
- Uncached input tokens. The genuinely new part of the request, at the normal input rate.
Add those up and you get your total input for the call.
Why the number is worth watching
Cache tokens are a direct, honest signal of how well your caching is working. A healthy long session should show most of its input arriving as cache reads, because the big stable prefix is being reused over and over. If that number stays low, something is busting the cache: an unstable system prompt, reordered context, or content that changes near the front on every request.
Reading them off a response
The usage object splits the input into what was written to the cache and what was read back from it.
const u = response.usage
console.log(u.cache_read_input_tokens) // served from cache, billed cheaper
console.log(u.cache_creation_input_tokens) // written to the cache this requestRelated terms
Prefix cache
A prefix cache lets a provider reuse the unchanged front of your request instead of reprocessing it, so repeated prefixes are cheaper and faster. It is the main reason keeping the start of your prompt stable pays off.
Read definition →ConceptInput tokens
Input tokens are the tokens you send in a request: the system prompt, the conversation history, loaded files, and tool definitions. You are billed for them, and they count against the context window.
Read definition →Explore it visually
- Attention variants: MHA, GQA, MQA, MLA and sliding windowsPress play and watch a KV-cache warehouse fill an H100: full attention needs 68.7 GB at Llama 3.1 8B’s full context, grouped-query attention cuts it to 17.2 GB, one shared head to a thirty-second, DeepSeek-V3’s latent attention compresses 71 times, and a sliding window caps it at 537 MB, all computed from the models’ published configs.
- KV cache and pagingServing a language model is a memory problem. See why reserving the declared context length for every request wastes most of your accelerator, and what paging the cache recovers.
- Prompt Caching VisualisedChange a prompt and trace exactly which cached blocks survive. Explore shared prefixes, block size, expiry and cache isolation in an interactive 2D and 3D prefix tree.
- Prompt caching: reuse the stable prefixA support bot resends the same 5,897-token prefix on every call. Watch real Claude Sonnet 5 calls read it back from a prompt cache for a tenth of the price, see a question-first prompt reuse nothing, and see one timestamp at the top of the system prompt throw the whole cache away.