Training is where a model comes from. You take an untrained network, feed it a vast corpus of text and code, and repeatedly nudge its parameters so its next-token predictions get less wrong. Do that at enormous scale and the model absorbs grammar, facts, coding patterns, and a surprising amount of reasoning ability. The output is the finished set of weights.
A one-time, frozen event
The two things worth remembering about training:
- It is expensive and rare. Training a frontier model costs a fortune in compute and happens on a schedule set by the provider, not by you. You consume the result via inference.
- It freezes knowledge in time. Whatever the model learned is fixed at the moment training ended. It has no awareness of anything that happened afterwards, which is the root of the knowledge-cutoff problem and a common source of confident-but-outdated answers.
Training vs. giving context
This is the single most useful implication for daily work: you do not "train" a coding agent by talking to it. Correcting it in a conversation changes nothing about its parameters. What you are actually doing is adjusting its context, the text in front of it right now. If you want it to know your codebase or your conventions, you supply that as context on each request, because the training door is closed.
Related terms
Model
A model is the trained artifact at the centre of every AI coding tool: a large file of numbers (parameters) that, given some text, produces the most likely continuation. When people say "which model are you using," this is the thing they mean.
Read definition →ConceptInference
Inference is the act of running a trained model to get an answer: text goes in, a prediction comes out. Every message you send to a coding agent is an inference. It is the opposite end of the lifecycle from training.
Read definition →AntipatternProvenHallucination
A hallucination is a confident, plausible-sounding output that is simply wrong: an invented API, a fabricated file path, a made-up citation. It is not the model lying. It is the model doing exactly what it always does, predicting plausible text, with no built-in sense of truth.
Read definition →Explore it visually
- Activation functions and normalisationPress play and watch ReLU rise, get sanded into GELU, become the SwiGLU valve, then see the outputs normalised, dropped out and added to the residual stream. Every slider still works mid-film.
- Backpropagation and vanishing gradientsPress play and watch backpropagation run the error back through one neuron by the chain rule, then through twelve layers, where the sigmoid lets it fade to 18 million times weaker and ReLU, careful initialisation and residual connections keep it alive.
- Diffusion models: from noise to imagePress play and watch pure static become a picture in fifty denoising steps: how training noises real pictures, why the cosine schedule beats the linear one, what the U-Net or DiT actually predicts, how a prompt steers sampling through classifier-free guidance, why too much guidance backfires, and why real systems diffuse a compressed latent, with every number computed from an exact toy model.
- Distributed training: data, tensor and pipeline parallelismA 7B model does not fit on one 80 GB card for training, and a 70B does not fit on eight. Slice the model four different ways across a rack and watch what each strategy costs in memory, traffic and idle time.
- Gradient accumulation: count every exampleTrace a tiny network, add gradients across microbatches of different sizes, expose the mean-of-means trap and make one correctly weighted update.
- LoRA: learn a correction, keep the baseFollow a frozen weight matrix and two trainable factors through a forward pass, one real training step, a rank limit, an exact parameter budget and a merge.
- Prompting, RAG and fine-tuning: what changes where?Prompting changes the instructions, RAG changes the evidence in the context, and fine-tuning changes the weights. One real support request runs through all three lanes, each fix repairs only its own fault, and the lane that learns from past replies also learns their out-of-date refund rule.
- RLHF to GRPO: reward models and the KL leashWatch one real GRPO step on a small model: eight answers to one maths question, a judge model standing in for the reward model, advantages computed against the group instead of a value model, and the reward hacking that followed within five steps, even with the KL penalty on.
- Scaling lawsSlide a compute budget across the Chinchilla map and watch the optimal split between parameters and tokens follow the ridge, with the tokens-per-parameter ratio sitting in the teens: close to, not exactly, the famous twenty.
- The loss landscapeDrop a ball anywhere on a loss surface and watch gradient descent live. Feel a learning rate diverge, fall into a decoy minimum, then add momentum and escape it.
- The optimiser raceSGD, momentum and Adam released together on an ill-conditioned valley. Drag the condition number and watch the exact learning rate at which SGD must explode, while Adam barely notices.
- What are embeddings?An embedding is the list of numbers a language model uses in place of a word. Watch a token id open a drawer in the embedding table, see similar words share similar numbers and fall into groups in 3D, measure the angle between two words, and do sums with meaning.