An automated check is a machine-verifiable gate that agent output has to pass: the test suite, the type checker, the linter, the build. It either passes or it fails, with no opinion and no negotiation. That objectivity is exactly what makes it valuable.
The backbone of trusting an agent
An agent writes plausible code that may or may not work. Automated checks are how you find out without reading all of it. A green type check proves the types line up; a passing test proves the behaviour holds. None of that depends on the code looking right, which is the whole point, because looking right is precisely what a model is good at faking.
This matters most when you are running AFK. With no human watching each step, checks are the guardrail. An agent that must make the tests pass before it is done has a hard, honest signal to work against, and it will iterate against that signal far more reliably than against your vague sense of quality.
- Fast and deterministic: same input, same verdict, every time.
- They catch regressions the agent cannot talk its way around.
- They are only as good as your coverage. Untested code has no gate.
The usual gates
These are the machine-verifiable gates an agent has to clear before its work counts as done. If they are green, most of the obvious failures are already ruled out.
npm run typecheck # tsc --noEmit, no type errors
npm test # unit tests pass
npm run lint # style and obvious bugs
npm run build # it actually compilesRelated terms
Automated review
Automated review is putting a change through an AI reviewer before a person sees it, so a second agent flags likely bugs, missed edge cases, and smells. It catches the obvious cheaply, but it does not replace human judgment about whether the change is right.
Read definition →PatternValidatedAFK
AFK means running an agent unattended for long stretches while you are away from the keyboard. It is only safe with strong guardrails and automated checks, since no human is watching each step.
Read definition →PatternProvenHuman review
Human review is a person actually reading what an agent produced, understanding it, and taking responsibility for shipping it. It is the final quality gate that tests and automated review can support but never replace.
Read definition →Explore it visually
- Draft, critique, revise, checkOne outage update goes round a revision loop: a writer model drafts, a critic model scores it against a five-point rubric and returns JSON, your code re-checks what it can count, and the writer revises one target at a time. Every text and verdict is a real model output, including the revision that made things worse.
- Image editing with masks and referencesSend one real photo, a mask, a reference mug and an instruction through OpenAI image models, then diff every result against the original. Under gpt-image-1.5 a tight mask still let 7 to 9 percent of the rest of the photo change. The same edits on gpt-image-2.5 changed under a tenth of a percent, with or without a mask, but a loose mask came back as a solid black block.
- Prompt optimisation can overfitA prompt optimiser keeps whichever edit raises the score on a handful of examples, which quietly turns that score into a training score. Watch a real loop take eight support tickets from 40 to 100 per cent, then lift a curtain on twenty-four tickets it never saw, where a simpler prompt from two steps in does better.
- RLHF to GRPO: reward models and the KL leashWatch one real GRPO step on a small model: eight answers to one maths question, a judge model standing in for the reward model, advantages computed against the group instead of a value model, and the reward hacking that followed within five steps, even with the KL penalty on.