A spec is a written description of what you want built and why, handed to the agent before it starts. It does not have to be long. It has to be clear: the goal, the constraints, what "done" looks like, and any edge cases you already know about.
Why a spec beats a one-liner
"Add auth" leaves a thousand decisions to the model, and it will guess at every one, sometimes well, often not in the way you meant. A spec is context you provide up front so the agent solves your problem instead of a plausible nearby one. Vague input reliably produces vague output; specific input is the cheapest quality lever you have.
A workable spec usually states:
- The goal, in one or two plain sentences.
- Constraints: the stack, the patterns, what not to touch.
- Acceptance criteria: how you will know it worked.
- Known edge cases and how they should behave.
Spec, ticket, artifact
A spec is close kin to a ticket and a handoff artifact. All three are just written context that tells an agent what to do. The label matters less than the discipline of writing the thing down before you set the agent loose.
What one looks like
A spec does not need to be long. It needs to be unambiguous about the goal, the constraints, and how you will know it is done.
# Spec: rate-limit login
## Goal
Block more than 5 failed logins per IP per minute.
## Constraints
Use the existing Redis client. Return HTTP 429 with a Retry-After header.
## Done when
New tests in the auth suite pass and the happy path is unchanged.Related terms
Ticket
A ticket is a scoped unit of work carrying enough context to act on. Well-formed tickets are ideal agent inputs.
Read definition →ConceptHandoff artifact
A handoff artifact is the concrete document produced at a handoff, recording what is done, what is next, and the key decisions. The next session reads it to get up to speed fast.
Read definition →ConceptAgent
An agent is a language model wrapped in a loop that lets it call tools, read the results, and decide what to do next. The model supplies the judgement; the loop and the tools give it hands.
Read definition →Explore it visually
- Draft, critique, revise, checkOne outage update goes round a revision loop: a writer model drafts, a critic model scores it against a five-point rubric and returns JSON, your code re-checks what it can count, and the writer revises one target at a time. Every text and verdict is a real model output, including the revision that made things worse.
- Image editing with masks and referencesSend one real photo, a mask, a reference mug and an instruction through OpenAI image models, then diff every result against the original. Under gpt-image-1.5 a tight mask still let 7 to 9 percent of the rest of the photo change. The same edits on gpt-image-2.5 changed under a tenth of a percent, with or without a mask, but a loose mask came back as a solid black block.
- Multi-shot prompting and visual continuityGenerate four shots of one barista with gpt-image-1.5 four ways and let a strict judge check every cut. Prompts written one at a time matched 16 of 42 checks; a shared continuity sheet, the approved first frame as a reference and a state line per shot took it to 54 of 54, because the model carries nothing from one shot to the next.
- Prompt optimisation can overfitA prompt optimiser keeps whichever edit raises the score on a handful of examples, which quietly turns that score into a training score. Watch a real loop take eight support tickets from 40 to 100 per cent, then lift a curtain on twenty-four tickets it never saw, where a simpler prompt from two steps in does better.
- The anatomy of a good promptA prompt has six pieces: task, audience, context, constraints, an example and an output format. Watch one real request assembled plate by plate, with the model’s recorded reply to every version, and see which piece changes the answer most.