Every request to a model has a budget: the context window. It is the total number of tokens the model can look at in one go, covering both what you send in and what it writes back. Modern models advertise large windows, but "large" is not "unlimited," and treating it as unlimited is the single most common way agent sessions go wrong.
What lives in the window
For a coding agent, the window is shared by a lot of competing tenants:
- The system prompt and tool definitions that set up the agent.
- Your instructions and any project rules the agent loads.
- File contents, command output, and search results the agent has pulled in.
- The full back-and-forth of the conversation so far.
Every one of those takes space, and space is finite. When the window fills, something has to give, and that is where quality quietly degrades.
Why it is the resource you manage
Two failure modes follow directly from the ceiling:
- Overflow. Push past the limit and the oldest or least-relevant content gets dropped or compacted. If the thing that got dropped was the instruction that actually mattered, the agent will confidently do the wrong thing.
- Dilution. Even well within the limit, a window stuffed with marginally-relevant text makes it harder for the model to attend to the few lines that count. More context is not automatically better context.
The craft of working with an agent is largely the craft of curating this window: giving it the files that matter, clearing out what is done, and pointing it at sources instead of pasting everything in.
Keeping this window clean is a whole discipline, which is why the context-engineering section of this dictionary exists: memory systems, progressive disclosure, and handoffs are all techniques for getting maximum value from a fixed number of tokens.
Related terms
AI
In the coding-agent world, "AI" almost always means a large language model: a system that predicts the next chunk of text from everything it has been shown. It is not a mind and it is not a database. It is a very good pattern completer.
Read definition →ConceptAgent
An agent is a language model wrapped in a loop that lets it call tools, read the results, and decide what to do next. The model supplies the judgement; the loop and the tools give it hands.
Read definition →ConceptMCP (Model Context Protocol)
MCP is an open standard for connecting agents to tools and data. Instead of hard-coding an integration into every agent, you run an MCP server once and any MCP-aware agent can use it.
Read definition →Explore it visually
- Attention variants: MHA, GQA, MQA, MLA and sliding windowsPress play and watch a KV-cache warehouse fill an H100: full attention needs 68.7 GB at Llama 3.1 8B’s full context, grouped-query attention cuts it to 17.2 GB, one shared head to a thirty-second, DeepSeek-V3’s latent attention compresses 71 times, and a sliding window caps it at 537 MB, all computed from the models’ published configs.
- Byte pair encodingType anything and watch a working BPE tokeniser build it from characters upward, one merge at a time. Shrink the vocabulary and see the same text cost more tokens.
- Chunking VisualisedPlay a short animation of one sentence split into readable text chunks. The same words separate in 3D, showing why each cut changes the context a retrieval system can see.
- Conversation history versus persistent memoryA model keeps nothing between calls, so an application resends the conversation and fills a context window. Follow one real conversation as the window overflows, truncation silently drops the turn that set a naming rule, and retrieved memory and pinned conventions bring it back, with every reply recorded.
- FlashAttentionFlashAttention is not a faster formula, it is a better memory schedule. Walk the tiles, watch the running softmax, and see why never writing the score matrix is worth re-reading the keys.
- GraphRAG VisualisedSee a Bristol launch question travel through a 3D graph into source passages and a cited answer. Adjust graph hops and passage budget, then press Play.
- KV cache and pagingServing a language model is a memory problem. See why reserving the declared context length for every request wastes most of your accelerator, and what paging the cache recovers.
- Lost in the middleA model accepting 128k tokens is not the same as a model using them. Sweep a fact through the context and watch retrieval collapse in the middle.
- Parallel work and the synthesis bottleneckSplit one research question across four parallel workers and merge their findings with a lead model, in real recorded runs. Parallel workers finish in a quarter of the time, but the merge becomes the slowest step, overlapping questions buy nothing, one slow branch holds everything up, and without a written merge rule a conflict between two sources quietly disappears.
- Prefill vs decode: time to first token and tokens per secondPress play and watch one GPU serve Llama 3.1 8B: prefill floods the whole prompt through at once and is limited by arithmetic, while decode drips one token at a time, rereads all 16 GB of weights for each, and is limited by memory bandwidth, with real timings measured on a live endpoint.
- Prompt caching: reuse the stable prefixA support bot resends the same 5,897-token prefix on every call. Watch real Claude Sonnet 5 calls read it back from a prompt cache for a tenth of the price, see a question-first prompt reuse nothing, and see one timestamp at the top of the system prompt throw the whole cache away.
- Rotary position embeddingsRoPE encodes position by rotating vectors, and the rotation cancels. Move a query and a key together and watch the attention score refuse to change.
- State space models vs attention: S4, Mamba and hybridsPress play and watch attention’s cache grow by 8 KiB with every token while a Mamba layer keeps one 128 KiB state, then follow the recurrence, the step size Δ, Mamba’s selectivity, the parallel scan and the exact recall a fixed state gives up, all computed from small, labelled toys.
- Vector search: HNSW, cosine similarity and BM25Follow one query down an HNSW tower as it hops from the roof to its ten nearest neighbours with 63 distance computations instead of 500, see what ef buys in 32 dimensions, then watch real embeddings and BM25 each miss a question the other gets right, and what reciprocal rank fusion does with the two.
- What are tokens?A language model never reads your words. A tokeniser cuts the text into pieces called tokens and swaps each one for an id number. Watch a machine do it to real sentences, with the real GPT-2 and GPT-4o splits and ids printed on every tile.