Models do not see letters or words. Before anything happens, your text is split into tokens: short chunks drawn from a fixed vocabulary. A common word like the is one token; a longer or rarer word like tokenization might be split into several. As a rough rule for English, one token is about four characters, and 100 tokens is roughly 75 words. Code, punctuation, and non-English text tokenize differently, often less efficiently.
Why you should care about a low-level detail
Tokens are not just an implementation detail. They are the unit almost everything is measured in:
- The [context window](/ai-coding-dictionary/context-window) is counted in tokens. A "200K context" means 200,000 tokens, not words or lines.
- Pricing is per token. You pay for tokens in and tokens out, so a verbose prompt or a giant pasted file has a direct cost.
- Limits bite in tokens. When an agent says it is running low on room, it is running low on tokens.
A practical consequence
Because tokenization is uneven, "small" inputs can be surprisingly expensive. A minified bundle, a wall of JSON, or a base64 blob can burn far more tokens than its character count suggests, while crowding out the code you actually want the model to focus on. Being deliberate about what you hand a model, and in what form, is really an exercise in spending tokens well.
Counting them
You do not have to guess. The provider can count the tokens in a request before you ever send it.
const { input_tokens } = await client.messages.countTokens({
model: 'claude-sonnet-5',
messages: [{ role: 'user', content: largePastedFile }],
})
console.log(input_tokens, 'tokens') // decide whether it is worth sendingRelated terms
Context window
The context window is the maximum amount of text, measured in tokens, that a model can consider for a single request. It is a hard ceiling, and it is the main resource you manage when working with an agent.
Read definition →ConceptModel
A model is the trained artifact at the centre of every AI coding tool: a large file of numbers (parameters) that, given some text, produces the most likely continuation. When people say "which model are you using," this is the thing they mean.
Read definition →ConceptInference
Inference is the act of running a trained model to get an answer: text goes in, a prediction comes out. Every message you send to a coding agent is an inference. It is the opposite end of the lifecycle from training.
Read definition →Explore it visually
- Attention as geometryAttention is dot products and a softmax. Drag the query and key vectors with your hands and feel it: alignment wins weight, and vector length is secretly a temperature.
- Attention head fingerprintsHeads are not redundant copies of each other. Step through twelve fingerprints and watch a previous-token head, an induction head and an attention sink behave in visibly different ways.
- Beam searchPlay beam search over a token tree built to trap greedy decoding. The most probable first word is not on the best sentence, and width two is enough to prove it.
- Byte pair encodingType anything and watch a working BPE tokeniser build it from characters upward, one merge at a time. Shrink the vocabulary and see the same text cost more tokens.
- Embedding spaceWords become points in a 3D space where nearness means similar meaning. Watch the clusters form, trace king to its nearest neighbours, see directions carry meaning, and find out what any flat picture of an embedding hides.
- KV cache and pagingServing a language model is a memory problem. See why reserving the declared context length for every request wastes most of your accelerator, and what paging the cache recovers.
- Mixture-of-experts routingHow a model can have a trillion parameters and run like a much smaller one. Route tokens to experts, watch the router collapse onto favourites, and see tokens dropped at capacity.
- Rotary position embeddingsRoPE encodes position by rotating vectors, and the rotation cancels. Move a query and a key together and watch the attention score refuse to change.
- Scaling lawsSlide a compute budget across the Chinchilla map and watch the optimal split between parameters and tokens follow the ridge, with the tokens-per-parameter ratio sitting in the teens: close to, not exactly, the famous twenty.
- Softmax, temperature and cross-entropyThe softmax function turns raw logits into probabilities in two moves: exponentiate, then divide by the sum. Watch a machine do it with the real numbers printed at every step, then heat it with temperature and read the loss off a gauge.
- The transformer block, explodedPull one transformer block apart like a watch movement and follow a token through it: embedding, attention, the residual stream, the MLP, and out to a distribution over the next word. Every part is counted, so you can see where the parameters actually live.
- Token decodingA model does not produce an answer, it produces a distribution over next tokens. Move temperature, top-k and top-p and watch the fan of possible continuations widen or collapse.
- What are embeddings?An embedding is the list of numbers a language model uses in place of a word. Watch a token id open a drawer in the embedding table, see similar words share similar numbers and fall into groups in 3D, measure the angle between two words, and do sums with meaning.
- What are tokens?A language model never reads your words. A tokeniser cuts the text into pieces called tokens and swaps each one for an id number. Watch a machine do it to real sentences, with the real GPT-2 and GPT-4o splits and ids printed on every tile.