Because a model learns nothing from your conversation, context is everything. It is the complete bundle of text handed to the model on a given request, and it is the model's entire view of your task. If a fact is not in the context, the model does not know it, no matter how obvious it seems to you.
What goes into it
For a coding agent, the context on any request is typically assembled from:
- The system prompt and the tool definitions.
- Your instructions and any project rules the agent loads.
- The back-and-forth of the current conversation.
- File contents, command output, and search results the agent has gathered.
All of that competes for the same fixed space, the context window, and it is all measured in tokens.
Context is the lever you pull
Almost everything you do to get better results from an agent is a form of context management. Pointing it at the right files, clearing out finished work, stating a constraint plainly, giving it an example: these are all ways of shaping what the model sees. When an agent gets something wrong, the first question is rarely "is the model bad" and almost always "did it have the right context."
Related terms
Context window
The context window is the maximum amount of text, measured in tokens, that a model can consider for a single request. It is a hard ceiling, and it is the main resource you manage when working with an agent.
Read definition →ConceptSystem prompt
The system prompt is the standing instruction placed at the very start of the context that sets the model’s role, rules, and tone before the conversation begins. It shapes every reply without being part of the back-and-forth.
Read definition →ConceptToken
A token is the unit of text a model reads and writes: a chunk that is usually part of a word, not a whole word or a single character. Everything is measured in tokens, including your context window and your bill.
Read definition →Explore it visually
- Chunking VisualisedPlay a short animation of one sentence split into readable text chunks. The same words separate in 3D, showing why each cut changes the context a retrieval system can see.
- Few-shot prompting: examples define the patternEight support tickets, one router, six prompts: no examples, useful examples, repetitive ones, one contradiction, one more example, and the house rules written down. Every landing is a real model answer recorded ten times, so you can see which examples resolve ambiguity and which only cost tokens.
- Lost in the middleA model accepting 128k tokens is not the same as a model using them. Sweep a fact through the context and watch retrieval collapse in the middle.
- Prompting, RAG and fine-tuning: what changes where?Prompting changes the instructions, RAG changes the evidence in the context, and fine-tuning changes the weights. One real support request runs through all three lanes, each fix repairs only its own fault, and the lane that learns from past replies also learns their out-of-date refund rule.
- RAG: from question to evidence to answerFollow one question through retrieval augmented generation: every handbook passage scored by real embedding similarity, the top three stacked into the prompt, and a recorded answer checked claim by claim. Then delete the right page, or add a stale one, and see retrieval and generation fail at different stages.
- The anatomy of a good promptA prompt has six pieces: task, audience, context, constraints, an example and an output format. Watch one real request assembled plate by plate, with the model’s recorded reply to every version, and see which piece changes the answer most.
- Vector search: HNSW, cosine similarity and BM25Follow one query down an HNSW tower as it hops from the roof to its ten nearest neighbours with 63 distance computations instead of 500, see what ef buys in 32 dimensions, then watch real embeddings and BM25 each miss a question the other gets right, and what reciprocal rank fusion does with the two.