Prompt → contract → verification: how we make agent work checkable
Write the done-criterion first, decide what a machine can check versus what needs a human, and keep evidence receipts so agent work can be trusted.
JPECOM Team4 min read
A prompt is a wish. A contract is a wish with edges, and verification is the part where somebody checks the edges. Most of the trouble people have with coding agents is not that the model is weak. It is that the work was never defined precisely enough to be checked, so "looks good" became the only acceptance test. This post describes the three-step habit we use to turn a request into something we can verify without reading every line.
Start with the done-criterion
Before writing the prompt, write down how you will know the work is finished. If you cannot, the task is not ready. This sounds trivial and it is the step most often skipped.
A good done-criterion has three properties:
- Observable. Someone other than the author can look and decide.
- Binary. It passes or it does not. "Better" is not a criterion; "the command exits zero" is.
- Independent of the agent's own report. It does not rely on the agent saying it worked.
Examples of weak versus strong criteria:
- Weak: "The login page should work."
- Strong: "The test command for the login module exits zero, and a new test covers a wrong-password case."
- Weak: "Clean up the utility functions."
- Strong: "No public function signature changes, the existing tests still pass, and the file has no function longer than a set limit."
Writing this first also improves the prompt. Once the criterion is clear, most of the instructions follow from it.
The contract: five parts
We write every agent task as a short contract. It is not long, but it has fixed parts:
- Goal. One or two sentences about the outcome, not the method.
- Interface. Inputs, outputs, names and shapes of anything other code depends on, written out in full. If the agent has to guess a name, it will guess differently from the next agent.
- Allowed files. What may be created or changed. Everything else is off-limits.
- Check command. The exact command that proves the work, and what its success looks like.
- Prohibitions. The things that must not happen: no new dependencies, no edits to existing tests, no network calls, and so on.
Goal: add a function that normalises phone numbers.
Interface: normalise_phone(raw: str) -> str | None
Allowed files: src/phone.py, tests/test_phone.py
Check: pytest tests/test_phone.py (must exit 0)
Do not: edit other files, add dependencies.
A contract of this size takes a couple of minutes to write and saves hours of back-and-forth. It also makes the task portable: any agent, or any human, can pick it up.
Verification: machine first, human second
Not everything can be checked by a program, so we split verification into layers and put the cheap ones first.
Layer 1: machine checks
Tests, type checks, linters, build commands, schema validators, and simple scripts that count things. These cost almost nothing, run in seconds and never get tired. Always run them first. If a machine check fails, there is no point in spending anyone's attention on the work.
Make the machine check strict enough to matter. A test suite that passes because the agent deleted the failing test is not evidence. Prohibiting edits to existing tests, and reviewing the diff for removed assertions, closes that door.
Layer 2: a second reader
For work that machines cannot judge, such as design quality, naming or whether the change matches the intent, a second reader reviews. We prefer that reader to be a different agent or person from the author. A model that wrote the code is a poor judge of it, for the same reason a writer is a poor proofreader of their own text. The reviewer gets the contract and the diff, and nothing else, so it judges against the stated criterion rather than the author's story.
Layer 3: a human for what matters
Anything touching money, customers, security or irreversible actions gets a person. The earlier layers exist to make that review short: by the time a human looks, the obvious failures are gone and the question is a narrow one.
Dirty-data testing
The check that catches the most is one that agents do not run by default: feed the work something unpleasant. Empty input, a very long string, the wrong type, a duplicate, text with unusual characters. If a function was only ever tested with the clean example from the prompt, you do not know how it behaves. We make "try to break it" an explicit step of review, because the happy path is what the agent already tested.
Evidence receipts
Verification is only useful if it leaves a trace. We ask for a short receipt with each completed task:
- The check command that was run and its exit status.
- The list of files that changed.
- Where to look: a path, a commit, a screenshot, an output file.
- Anything that was skipped or could not be checked, stated plainly.
The last point is the most valuable. A report that says "everything was checked" is less trustworthy than one that says "these four things were checked and these two were not". Honest gaps are information. Hidden gaps are future incidents.
The orchestrator, not the agent, should produce the receipt wherever possible. If the agent writes its own receipt, you are back to trusting a narrative.
When the criterion cannot be written
Sometimes you truly cannot define done up front, for example in exploratory work. In that case, shrink the task until you can: make the deliverable a written plan, a prototype, or a list of options, and define done for that. Then the next task, which builds on the chosen option, gets a real contract.
Takeaways
- Write the done-criterion before the prompt: observable, binary, independent of the agent.
- Use a five-part contract: goal, interface, allowed files, check command, prohibitions.
- Verify in layers: machine checks, then a different reader, then a human for high-stakes work.
- Test with unpleasant input, not just the example from the prompt.
- Keep an evidence receipt, including what was not checked.
Related posts
Multi-agent orchestration: the mistakes that cost us weeks
Agents that say done without evidence, retry loops, file ownership clashes, and our rules: two failures means stop, measure output not activity.
5 min read
Why we run coding agents on subscriptions, not API keys
Predictable cost, quota as a budget, and the cases where pay-per-token API access is still the right call for coding agents.
5 min read
AI Update, October 2026: Notable Shifts for People Who Build Products
Trends we have been seeing: CLI agents are maturing, subscriptions come in more tiers, local models are worth a look, and verification tooling is moving to the center.
4 min read