Every model provider request has two sides. Input tokens are the tokens you send in: the system prompt, the whole conversation so far, any files the agent has read, the tool definitions, and your latest message. They are also called prompt tokens. The provider adds them all up, and that total is what the model reads before it writes a single word back.
What counts as input
It is easy to underestimate this, because most input is not typed by you. On a real coding request the input is dominated by:
- The system prompt and every tool definition, sent on every request.
- The full history of the session, which only grows.
- File contents and command output the agent has pulled in.
Your actual instruction is often the smallest part.
Why input tokens deserve attention
Two reasons, one about cost and one about quality:
- You pay for them. Input tokens are billed, usually at a lower rate than output tokens, but there are far more of them, so they often dominate the bill on a long session.
- They fill the [context window](/ai-coding-dictionary/context-window). Input is what consumes the window, and a window packed with marginally relevant input both costs more and dilutes the model's attention.
Where you see them
After any request, the usage object reports exactly how many input tokens it cost.
const response = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 1024,
messages,
})
console.log(response.usage.input_tokens) // everything you sent, countedRelated terms
Token
A token is the unit of text a model reads and writes: a chunk that is usually part of a word, not a whole word or a single character. Everything is measured in tokens, including your context window and your bill.
Read definition →ConceptOutput tokens
Output tokens are the tokens a model generates in its response, including any hidden reasoning. They are usually priced higher than input tokens, and turning up effort produces more of them.
Read definition →ConceptContext window
The context window is the maximum amount of text, measured in tokens, that a model can consider for a single request. It is a hard ceiling, and it is the main resource you manage when working with an agent.
Read definition →Explore it visually
- Byte pair encodingType anything and watch a working BPE tokeniser build it from characters upward, one merge at a time. Shrink the vocabulary and see the same text cost more tokens.
- Few-shot prompting: examples define the patternEight support tickets, one router, six prompts: no examples, useful examples, repetitive ones, one contradiction, one more example, and the house rules written down. Every landing is a real model answer recorded ten times, so you can see which examples resolve ambiguity and which only cost tokens.
- Prefill vs decode: time to first token and tokens per secondPress play and watch one GPU serve Llama 3.1 8B: prefill floods the whole prompt through at once and is limited by arithmetic, while decode drips one token at a time, rereads all 16 GB of weights for each, and is limited by memory bandwidth, with real timings measured on a live endpoint.
- Prompt Caching VisualisedChange a prompt and trace exactly which cached blocks survive. Explore shared prefixes, block size, expiry and cache isolation in an interactive 2D and 3D prefix tree.
- Prompt caching: reuse the stable prefixA support bot resends the same 5,897-token prefix on every call. Watch real Claude Sonnet 5 calls read it back from a prompt cache for a tenth of the price, see a question-first prompt reuse nothing, and see one timestamp at the top of the system prompt throw the whole cache away.
- What are tokens?A language model never reads your words. A tokeniser cuts the text into pieces called tokens and swaps each one for an id number. Watch a machine do it to real sentences, with the real GPT-2 and GPT-4o splits and ids printed on every tile.