Send a model the exact same prompt twice and you can get two different answers. That is non-determinism, and it is by design, not a bug. During next-token prediction the model produces a probability distribution over possible next tokens, and then it usually samples from that distribution rather than always taking the single most likely token. A setting called temperature controls how much randomness gets mixed in. Higher temperature, more variety; lower, more repetition.
Why providers do this
A little randomness makes output feel less robotic and helps the model escape repetitive ruts. The trade is reproducibility. Even at very low temperature you are not guaranteed identical results, because inference runs on batched hardware where tiny numerical variations creep in.
What it means for coding work
This is easy to forget until it bites you:
- A passing run is not proof. An agent solving a task once does not mean it will solve it every time. If reliability matters, run it more than once.
- Do not hard-code on exact wording. Tests or scripts that assume the model returns a specific string will be flaky. Assert on behaviour or structure instead.
- Bugs can be intermittent. A prompt that fails one time in five is still broken. Chase the pattern, not the single lucky success.
Turning it down
Lowering the temperature reduces variation. Zero is the most deterministic setting, though even then output is not guaranteed to be byte-identical.
const response = await client.messages.create({
model: 'claude-sonnet-5',
max_tokens: 1024,
temperature: 0,
messages,
})Related terms
Inference
Inference is the act of running a trained model to get an answer: text goes in, a prediction comes out. Every message you send to a coding agent is an inference. It is the opposite end of the lifecycle from training.
Read definition →ConceptNext-token prediction
Next-token prediction is the one job a language model does: given the text so far, predict the most likely next token, add it, and repeat. It is both the training objective and what runs at inference.
Read definition →ConceptEffort
Effort is a dial for how much internal reasoning a model spends before it answers. Turn it up for genuinely hard problems; you pay for it in latency and extra output tokens.
Read definition →Explore it visually
- Few-shot prompting: examples define the patternEight support tickets, one router, six prompts: no examples, useful examples, repetitive ones, one contradiction, one more example, and the house rules written down. Every landing is a real model answer recorded ten times, so you can see which examples resolve ambiguity and which only cost tokens.
- Image editing with masks and referencesSend one real photo, a mask, a reference mug and an instruction through OpenAI image models, then diff every result against the original. Under gpt-image-1.5 a tight mask still let 7 to 9 percent of the rest of the photo change. The same edits on gpt-image-2.5 changed under a tenth of a percent, with or without a mask, but a loose mask came back as a solid black block.
- Multi-shot prompting and visual continuityGenerate four shots of one barista with gpt-image-1.5 four ways and let a strict judge check every cut. Prompts written one at a time matched 16 of 42 checks; a shared continuity sheet, the approved first frame as a reference and a state line per shot took it to 54 of 54, because the model carries nothing from one shot to the next.
- Prompt optimisation can overfitA prompt optimiser keeps whichever edit raises the score on a handful of examples, which quietly turns that score into a training score. Watch a real loop take eight support tickets from 40 to 100 per cent, then lift a curtain on twenty-four tickets it never saw, where a simpler prompt from two steps in does better.
- RLHF to GRPO: reward models and the KL leashWatch one real GRPO step on a small model: eight answers to one maths question, a judge model standing in for the reward model, advantages computed against the group instead of a value model, and the reward hacking that followed within five steps, even with the KL penalty on.
- Speculative decodingRun a second model and finish sooner. Watch a cheap draft model propose tokens, a large model verify them in a single batched pass, and the speedup curve peak and then fall.
- Token decodingA model does not produce an answer, it produces a distribution over next tokens. Move temperature, top-k and top-p and watch the fan of possible continuations widen or collapse.