Chapter 9 · The Agent Playbook for Software Engineers
Verification
An agent stops when the work looks done. If nothing else can tell it otherwise, "looks done" is the only signal in the system, and you are the one holding it.
Here is the sequence almost everyone runs into in their first month with a CLI coding agent. You describe a task. The agent reads some files, writes some code, and reports that it has finished. You run the test suite. Two tests fail. You paste the failures back in. The agent apologises, changes something, and reports that it has finished. You run the test suite again.
Notice what is actually happening in that loop: you are the verification step. Every mistake the agent makes waits patiently for you to notice it. The agent is not slow and it is not incapable. It simply has no way to find out whether it is right, so it uses the only available proxy, which is whether the output looks like the kind of thing that would be correct.
The fix is one sentence long and it is the highest-leverage habit in this entire book:
Give the agent a check it can run.
A test suite. A build exit code. A linter. A script that diffs output against a fixture. A browser screenshot compared against a design. Anything that produces a pass or a fail without you in the room. Once the agent can run the check itself, the loop closes on its own, and the difference is not incremental: it is the difference between a session you have to watch and a session you can walk away from.
What counts as a check
A valid check has three properties. It is runnable by the agent with the permissions it already has. It produces an unambiguous pass or fail rather than a paragraph of prose. And it is cheap enough to run repeatedly, because the whole point is that it runs after every attempt rather than once at the end.
In descending order of how often they turn out to be the right tool:
| Check | Good for | Watch out for |
|---|---|---|
| A single test file | Bug fixes, small features, anything with a reproduction | Running the whole suite when one file would do. It is slower and it buries the signal. |
| Typecheck or build exit code | Refactors, renames, migrations across many files | A green build is not a correct build. Pair it with something behavioural. |
| A fixture diff script | Output formats, codegen, data transforms | The agent editing the fixture instead of the code. Say explicitly that the fixture is frozen. |
| A screenshot compared to a reference | UI work | Needs an explicit "take a screenshot, compare, list differences" instruction. Agents do not do this unprompted. |
| A linter or formatter | Style conformance only | Never sufficient on its own. It says the code is tidy, not that it works. |
Notice what is missing from that list: "read the code and confirm it is right". Asking an agent to grade its own reasoning is not a check. It is another opinion from the same source that produced the work, and it will usually agree with itself.
Putting the check in the prompt
The cheapest version of this works today, on any task, with no setup at all: name the check in the same message as the task, and tell the agent to iterate until it passes.
Weak
implement a function that validates email addresses
Strong
write a validateEmail function. example test cases: user@example.com is true,
invalid is false, user@.com is false. run the tests after implementing
Three words are doing most of the work there: "run the tests". Without them the agent
writes a function and stops. With them, the same session produces a function, a test run,
and, more often than you would expect, a second revision that the agent made on its own
because the first one failed on user@.com.
The pattern generalises to work that is not a unit test.
Weak
make the dashboard look better
Strong
[paste screenshot] implement this design. take a screenshot of the result and
compare it to the original. list differences and fix them
Weak
the build is failing
Strong
the build fails with this error: [paste error]. fix it and verify the build
succeeds. address the root cause, don't suppress the error
That last clause matters more than it looks. An agent asked to make a build pass has a legal move available that you did not intend: delete the failing assertion, catch and swallow the exception, or add the file to an ignore list. All of those produce a passing check. Naming the root cause as the requirement closes the loophole, and it is the reason "address the root cause, don't suppress the error" appears verbatim in so many good prompts.
Ask for the evidence
Tell the agent to show the test output, the command it ran and what it returned, or the screenshot. Reading two lines of pytest output is faster than re-running the verification yourself, and it catches the case where the agent believes it ran a check that it never actually ran.
Four levels of gating
Putting the check in the prompt is level one. There are three more, and each trades setup effort for attention you get back later. You do not need all four. You need to know which one a given piece of work deserves.
- In one prompt. Ask the agent to run the check and iterate within the same message. Zero setup, works on any task today, and it is what most sessions should use. The weakness is that the agent decides when it is satisfied.
- Across a session. Set the check as a persistent goal condition. A separate evaluator re-checks it after every turn and the agent keeps working until it holds. This survives the agent losing track of the requirement twenty turns later, which is the most common failure in a long session.
- As a deterministic gate. A stop hook runs your check as a script and blocks the turn from ending until it passes. This is the strongest form: it is your code, not the model's judgement, that decides whether the work is done. Note that the tool will override the hook and end the turn after a run of consecutive blocks, so a hook that can never pass produces a stuck session rather than an infinite one.
- By a second opinion. A verification subagent, running in a fresh context, reviews the diff against the criteria. The point is that the agent doing the work is not the one grading it.
The practical rule: level one for anything you are watching, level three for anything you intend to leave running. Levels two and four are for work where correctness matters more than iteration speed.
A verification loop the agent can drive itself
#!/usr/bin/env bash set -euo pipefail pnpm lint pnpm typecheck pnpm test --run
Four lines, checked into the repository, named in the project memory file. That script is the single highest-return artifact in most codebases that use agents, because it converts "is this done?" from a judgement call into an exit code, and because every future session inherits it for free.
The adversarial review step
The longer an agent works unattended, the more an independent check matters before you count the work as done. A reviewer running in a fresh subagent context sees only the diff and the criteria you give it. It does not see the reasoning that produced the change, so it evaluates the result on its own terms rather than being carried along by the argument that justified it.
Use a subagent to review the rate limiter diff against PLAN.md. Check that every requirement is implemented, the listed edge cases have tests, and nothing outside the task's scope changed. Report gaps, not style preferences.
Three constraints in that prompt are load-bearing. "Against PLAN.md" gives the reviewer a fixed standard instead of its own taste. "Nothing outside the task's scope changed" catches the most under-reported agent failure, which is unrequested edits to adjacent files. "Report gaps, not style preferences" is what stops the review from becoming a list of opinions about naming.
The counter-warning
A reviewer prompted to find gaps will usually report some, even when the work is sound, because that is what it was asked to do. Chasing every finding leads directly to over-engineering: extra abstraction layers, defensive code, and tests for cases that cannot happen. Tell the reviewer to flag only gaps that affect correctness or the stated requirements, and treat everything else as optional.
This is worth sitting with, because it is the failure mode that catches careful people. Adding a review step feels unambiguously good, so the instinct is to act on everything it returns. Two weeks later the codebase has three defensive wrappers around a function that only ever receives validated input, and none of it was requested by anyone. The review step is a filter, not a backlog.
What this looks like when it is working
A session with a real verification loop reads differently. The agent writes code, runs the check, sees a failure, reads the failure, and revises, all before it says anything to you. What arrives in your terminal is not "I have implemented the function" but "I implemented it, the first version failed on the empty-string case, here is the output, here is the fix, all seven tests now pass." You read four lines and a test summary instead of a diff.
The economics change too. A plan-then-execute loop with a check typically costs less than a single unguided attempt at the same task, because the expensive part of agent work is never the first attempt. It is the rework, and the human time spent discovering that rework is needed.
When there is no check available
Some work genuinely has no automatable check: exploratory refactors, documentation, judgement calls about architecture. For those, do not pretend. Keep the session short, keep the diff small, and read it. The mistake is not working without a check. The mistake is working without a check while behaving as though you have one.
What to remember
- Without a check it can run, "looks done" is the only signal the agent has, and you become the verification loop by default.
- Name the check in the prompt. "Run the tests after implementing" is three words and it changes the shape of the session.
- Say "address the root cause, don't suppress the error", because suppressing it is a legal move that satisfies a naive check.
- Ask for evidence rather than assertions. Test output beats "done".
- Escalate to a deterministic stop gate for anything you intend to leave unattended.
- Use a fresh-context reviewer for long unattended work, and tell it to report only gaps that affect correctness. Otherwise you will over-engineer in response to its opinions.
End of chapter 9
Eleven more chapters, built the same way.
Chapter five is the anatomy of a prompt that works: the four components that appear in every request you can merge, and what breaks when one is missing. Chapter seven is a library of fifty-one prompts organised by the job you are doing. Chapter four is on what belongs in a project memory file and what quietly makes it worse.
The full book is 104 pages across 12 chapters, plus a quick-reference appendix. It is a PDF. It is $79 once, with no subscription, and you can ask for a refund within 60 days without explaining yourself. After checkout we will offer you the ecommerce playbook once, at the bundle difference; declining changes nothing about your order.
Not buying today
This chapter has a shelf life, and we will tell you when it expires.
Permission flags get renamed. Hook behaviour changes. A paragraph that was accurate in July stops being accurate. When something in the chapter you just read goes out of date, we rewrite it and send you the correction. That is the entire list: corrections to this chapter, and a note when a new edition of the book ships. No sequence, no launch campaign, and you can leave from the first email.
Done. Check your inbox for a confirmation, which also carries the unsubscribe link.
That did not work. Check the address and try again, or email us and we will add you by hand.
You are unsubscribed. You will not hear from us about this again.