Output tokens are the other half of a request: the tokens the model actually generates. The reply you read, the code it writes, the tool calls it emits, and on reasoning models a chunk of hidden thinking too, all of it is output. They are sometimes called completion tokens.
Slower and pricier than input
Output behaves differently from input tokens in two ways that matter:
- They cost more. Providers almost always price output tokens higher than input, often several times higher. Generating text is the expensive part.
- They are produced one at a time. The model writes output sequentially, each token after the last, which is why a long answer streams in slowly and takes real wall-clock time. Input, by contrast, is read in one pass.
Where they quietly pile up
The surprise on modern models is how much output you cannot see. When you raise effort, the model reasons more before answering, and that reasoning is billed as output even if it never reaches your screen. A short final answer can sit on top of a large, invisible pile of reasoning tokens.
Because output is the slow, costly side of a request, brevity has value. Asking an agent to write a novel where a sentence would do is not just noise, it is latency and money.
Where you see them
The same usage object reports what the model generated. Output tokens are usually the pricier half of the bill.
const response = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 1024,
messages,
})
console.log(response.usage.output_tokens) // what the model wrote backRelated terms
Token
A token is the unit of text a model reads and writes: a chunk that is usually part of a word, not a whole word or a single character. Everything is measured in tokens, including your context window and your bill.
Read definition →ConceptInput tokens
Input tokens are the tokens you send in a request: the system prompt, the conversation history, loaded files, and tool definitions. You are billed for them, and they count against the context window.
Read definition →ConceptEffort
Effort is a dial for how much internal reasoning a model spends before it answers. Turn it up for genuinely hard problems; you pay for it in latency and extra output tokens.
Read definition →Explore it visually
- Byte pair encodingType anything and watch a working BPE tokeniser build it from characters upward, one merge at a time. Shrink the vocabulary and see the same text cost more tokens.
- Continuous batchingThe same twelve requests scheduled twice: static batches strand GPU slots behind their slowest member, continuous admission refills them next step. Watch the score diverge live.
- Prefill vs decode: time to first token and tokens per secondPress play and watch one GPU serve Llama 3.1 8B: prefill floods the whole prompt through at once and is limited by arithmetic, while decode drips one token at a time, rereads all 16 GB of weights for each, and is limited by memory bandwidth, with real timings measured on a live endpoint.
- Prompt caching: reuse the stable prefixA support bot resends the same 5,897-token prefix on every call. Watch real Claude Sonnet 5 calls read it back from a prompt cache for a tenth of the price, see a question-first prompt reuse nothing, and see one timestamp at the top of the system prompt throw the whole cache away.
- Speculative decodingRun a second model and finish sooner. Watch a cheap draft model propose tokens, a large model verify them in a single batched pass, and the speedup curve peak and then fall.
- The anatomy of a good promptA prompt has six pieces: task, audience, context, constraints, an example and an output format. Watch one real request assembled plate by plate, with the model’s recorded reply to every version, and see which piece changes the answer most.
- What are tokens?A language model never reads your words. A tokeniser cuts the text into pieces called tokens and swaps each one for an id number. Watch a machine do it to real sentences, with the real GPT-2 and GPT-4o splits and ids printed on every tile.