A model is, physically, a giant list of parameters: numbers, often billions of them, that were tuned during training. People also call them weights. They are not settings you configure. They are values the training process discovered on its own, and together they encode every pattern the model picked up, from grammar to how a React component is usually structured.
Size is a parameter count
When you hear "a 70-billion-parameter model," that number is the parameter count. More parameters give a model more capacity to store patterns, which usually means stronger reasoning over messy code, at the cost of more compute, higher latency, and a bigger bill. It is a rough proxy for capability, not a guarantee: how a model was trained matters just as much as how many parameters it has.
Frozen at inference
The key thing for daily work is that parameters are set once and then frozen. During inference, when you actually use the model, nothing about them changes. Your conversation does not nudge a single weight. That is why a coding agent cannot "learn" your codebase by chatting, and why everything it knows about your specific project has to arrive as context on each request.
Related terms
Model
A model is the trained artifact at the centre of every AI coding tool: a large file of numbers (parameters) that, given some text, produces the most likely continuation. When people say "which model are you using," this is the thing they mean.
Read definition →ConceptTraining
Training is the process that produces a model: showing it enormous amounts of text and adjusting its parameters until it gets good at predicting what comes next. It happens once, before you ever use the model.
Read definition →ConceptInference
Inference is the act of running a trained model to get an answer: text goes in, a prediction comes out. Every message you send to a coding agent is an inference. It is the opposite end of the lifecycle from training.
Read definition →Explore it visually
- Activation functions and normalisationPress play and watch ReLU rise, get sanded into GELU, become the SwiGLU valve, then see the outputs normalised, dropped out and added to the residual stream. Every slider still works mid-film.
- Backpropagation and vanishing gradientsPress play and watch backpropagation run the error back through one neuron by the chain rule, then through twelve layers, where the sigmoid lets it fade to 18 million times weaker and ReLU, careful initialisation and residual connections keep it alive.
- Diffusion models: from noise to imagePress play and watch pure static become a picture in fifty denoising steps: how training noises real pictures, why the cosine schedule beats the linear one, what the U-Net or DiT actually predicts, how a prompt steers sampling through classifier-free guidance, why too much guidance backfires, and why real systems diffuse a compressed latent, with every number computed from an exact toy model.
- Distributed training: data, tensor and pipeline parallelismA 7B model does not fit on one 80 GB card for training, and a 70B does not fit on eight. Slice the model four different ways across a rack and watch what each strategy costs in memory, traffic and idle time.
- Embedding spaceWords become points in a 3D space where nearness means similar meaning. Watch the clusters form, trace king to its nearest neighbours, see directions carry meaning, and find out what any flat picture of an embedding hides.
- Gradient accumulation: count every exampleTrace a tiny network, add gradients across microbatches of different sizes, expose the mean-of-means trap and make one correctly weighted update.
- LoRA: learn a correction, keep the baseFollow a frozen weight matrix and two trainable factors through a forward pass, one real training step, a rank limit, an exact parameter budget and a merge.
- Matrix multiplication on a GPUA GEMM is exactly 2MNK operations whatever runs it. Watch a GPU tile one, then compare three ways of building the hardware for it, CUDA cores, tensor cores and a TPU systolic array, and see why the fast ones are so hard to feed.
- Mixture-of-experts routingHow a model can have a trillion parameters and run like a much smaller one. Route tokens to experts, watch the router collapse onto favourites, and see tokens dropped at capacity.
- Prefill vs decode: time to first token and tokens per secondPress play and watch one GPU serve Llama 3.1 8B: prefill floods the whole prompt through at once and is limited by arithmetic, while decode drips one token at a time, rereads all 16 GB of weights for each, and is limited by memory bandwidth, with real timings measured on a live endpoint.
- Prompting, RAG and fine-tuning: what changes where?Prompting changes the instructions, RAG changes the evidence in the context, and fine-tuning changes the weights. One real support request runs through all three lanes, each fix repairs only its own fault, and the lane that learns from past replies also learns their out-of-date refund rule.
- Quantisation and outliersWhy int4 sometimes works and sometimes destroys a model. Watch a handful of outlier channels consume the entire numeric range, and watch per-channel scaling give it back.
- RLHF to GRPO: reward models and the KL leashWatch one real GRPO step on a small model: eight answers to one maths question, a judge model standing in for the reward model, advantages computed against the group instead of a value model, and the reward hacking that followed within five steps, even with the KL penalty on.
- Rotary position embeddingsRoPE encodes position by rotating vectors, and the rotation cancels. Move a query and a key together and watch the attention score refuse to change.
- Scaling lawsSlide a compute budget across the Chinchilla map and watch the optimal split between parameters and tokens follow the ridge, with the tokens-per-parameter ratio sitting in the teens: close to, not exactly, the famous twenty.
- Superposition and sparse decodingPack five feature directions into two coordinates, measure the resulting cross-talk, and run a sparse decoder that reveals both what it can recover and what compression has erased.
- The loss landscapeDrop a ball anywhere on a loss surface and watch gradient descent live. Feel a learning rate diverge, fall into a decoy minimum, then add momentum and escape it.
- The optimiser raceSGD, momentum and Adam released together on an ill-conditioned valley. Drag the condition number and watch the exact learning rate at which SGD must explode, while Adam barely notices.
- The transformer block, explodedPull one transformer block apart like a watch movement and follow a token through it: embedding, attention, the residual stream, the MLP, and out to a distribution over the next word. Every part is counted, so you can see where the parameters actually live.
- What are embeddings?An embedding is the list of numbers a language model uses in place of a word. Watch a token id open a drawer in the embedding table, see similar words share similar numbers and fall into groups in 3D, measure the angle between two words, and do sums with meaning.