A model is what training produces: a fixed set of parameters, often billions of them, stored in a file you load and run. Those parameters encode the patterns the model learned from its training data. Nothing about the model changes when you use it; you feed in text and it computes an output. All the "intelligence" is baked into those frozen numbers.
Parameters, not rules
The parameters are weights, learned automatically during training, not rules a human wrote. That is why you cannot open a model and find the line that decides how it writes a for-loop. The behaviour is distributed across the whole network. It also explains why models are hard to fully predict: you are working with learned statistics, not a program someone specified.
Why "which model" matters
Models differ in ways that directly affect your work:
- Size and capability. Larger, more capable models tend to reason better over long, messy code, at higher cost and latency.
- Training and recency. Two models trained differently will have different strengths and different knowledge cutoffs.
- Specialisation. Some models are tuned for chat, some for code, some for tool use.
When a coding tool lets you pick a model, you are trading off quality, speed, and cost. For a quick rename, a small fast model is fine. For untangling a subtle bug across ten files, reach for the strongest one you have.
Related terms
AI
In the coding-agent world, "AI" almost always means a large language model: a system that predicts the next chunk of text from everything it has been shown. It is not a mind and it is not a database. It is a very good pattern completer.
Read definition →ConceptToken
A token is the unit of text a model reads and writes: a chunk that is usually part of a word, not a whole word or a single character. Everything is measured in tokens, including your context window and your bill.
Read definition →ConceptInference
Inference is the act of running a trained model to get an answer: text goes in, a prediction comes out. Every message you send to a coding agent is an inference. It is the opposite end of the lifecycle from training.
Read definition →Explore it visually
- Activation functions and normalisationPress play and watch ReLU rise, get sanded into GELU, become the SwiGLU valve, then see the outputs normalised, dropped out and added to the residual stream. Every slider still works mid-film.
- Attention head fingerprintsHeads are not redundant copies of each other. Step through twelve fingerprints and watch a previous-token head, an induction head and an attention sink behave in visibly different ways.
- Backpropagation and vanishing gradientsPress play and watch backpropagation run the error back through one neuron by the chain rule, then through twelve layers, where the sigmoid lets it fade to 18 million times weaker and ReLU, careful initialisation and residual connections keep it alive.
- Diffusion models: from noise to imagePress play and watch pure static become a picture in fifty denoising steps: how training noises real pictures, why the cosine schedule beats the linear one, what the U-Net or DiT actually predicts, how a prompt steers sampling through classifier-free guidance, why too much guidance backfires, and why real systems diffuse a compressed latent, with every number computed from an exact toy model.
- Distributed training: data, tensor and pipeline parallelismA 7B model does not fit on one 80 GB card for training, and a 70B does not fit on eight. Slice the model four different ways across a rack and watch what each strategy costs in memory, traffic and idle time.
- Embedding spaceWords become points in a 3D space where nearness means similar meaning. Watch the clusters form, trace king to its nearest neighbours, see directions carry meaning, and find out what any flat picture of an embedding hides.
- Gradient accumulation: count every exampleTrace a tiny network, add gradients across microbatches of different sizes, expose the mean-of-means trap and make one correctly weighted update.
- LoRA: learn a correction, keep the baseFollow a frozen weight matrix and two trainable factors through a forward pass, one real training step, a rank limit, an exact parameter budget and a merge.
- Mixture-of-experts routingHow a model can have a trillion parameters and run like a much smaller one. Route tokens to experts, watch the router collapse onto favourites, and see tokens dropped at capacity.
- Quantisation and outliersWhy int4 sometimes works and sometimes destroys a model. Watch a handful of outlier channels consume the entire numeric range, and watch per-channel scaling give it back.
- RLHF to GRPO: reward models and the KL leashWatch one real GRPO step on a small model: eight answers to one maths question, a judge model standing in for the reward model, advantages computed against the group instead of a value model, and the reward hacking that followed within five steps, even with the KL penalty on.
- Scaling lawsSlide a compute budget across the Chinchilla map and watch the optimal split between parameters and tokens follow the ridge, with the tokens-per-parameter ratio sitting in the teens: close to, not exactly, the famous twenty.
- Speculative decodingRun a second model and finish sooner. Watch a cheap draft model propose tokens, a large model verify them in a single batched pass, and the speedup curve peak and then fall.
- State space models vs attention: S4, Mamba and hybridsPress play and watch attention’s cache grow by 8 KiB with every token while a Mamba layer keeps one 128 KiB state, then follow the recurrence, the step size Δ, Mamba’s selectivity, the parallel scan and the exact recall a fixed state gives up, all computed from small, labelled toys.
- Superposition and sparse decodingPack five feature directions into two coordinates, measure the resulting cross-talk, and run a sparse decoder that reveals both what it can recover and what compression has erased.
- The loss landscapeDrop a ball anywhere on a loss surface and watch gradient descent live. Feel a learning rate diverge, fall into a decoy minimum, then add momentum and escape it.
- The optimiser raceSGD, momentum and Adam released together on an ill-conditioned valley. Drag the condition number and watch the exact learning rate at which SGD must explode, while Adam barely notices.
- The transformer block, explodedPull one transformer block apart like a watch movement and follow a token through it: embedding, attention, the residual stream, the MLP, and out to a distribution over the next word. Every part is counted, so you can see where the parameters actually live.
- What are embeddings?An embedding is the list of numbers a language model uses in place of a word. Watch a token id open a drawer in the embedding table, see similar words share similar numbers and fall into groups in 3D, measure the angle between two words, and do sums with meaning.