❯ Reviewing agent-written code like it matters

by Jonas Reyes — the builder's desk, Side Quest Studios September 29, 2026

That pull request your coding agent just opened is not a finished artifact — it's a draft that happens to be confidently written. Here's the review checklist that keeps agent code from becoming next month's incident.

← back to articles

The one sentence

Before you merge anything written by an AI agent, read the diff, run the tests, and make the agent prove the failure case it claims to handle.

Why this matters

Coding agents are fluent. They produce compilable, well-named, plausible code at speed — and that is exactly the danger. A human who writes a wrong API call usually finds it while testing. An agent that believes the API works a certain way will write a test that passes on a stub and ships the wrong behavior. Reviewing agent output with the same rigor you'd apply to a junior dev's first PR is what separates "fast" from "on fire."

The failure has a signature: the test and the code encode the same wrong assumption, so they confirm each other and confirm nothing about reality. A green suite is evidence that the suite ran, not that the software works.

The checklist

1. Diff first, summary second. The agent's own summary is marketing. Look at the actual changes. A pull request with 400 lines of "refactor" and a 12-line feature buried inside is a red flag, not a convenience.

you@project:~ bash (80x24)
you@project:~$ gh pr diff 142 | wc -l
487

2. Run the real test suite, in the real environment. Not the mocked one the agent used. pytest -q on your machine, or the CI it will actually ship through. A green unit test in isolation is not the same as a green integration.

you@project:~ bash (80x24)
you@project:~$ pytest -q tests/integration
...............
15 passed, 1 skipped in 3.1s

3. Hunt the negative case. Agents write the happy path beautifully. Ask them — or look — for what happens when the input is empty, the network times out, or the row doesn't exist. If the error handling is "it'll be fine," it won't be.

A retry wrapper is the classic. I reviewed an agent's send_with_retry() last week: fifteen clean lines and a passing test named test_retries_on_timeout. The test mocked the first call to raise Timeout and asserted the function was invoked twice. Green in isolation, green in CI. In production it double-sent every welcome email whose SMTP send had actually succeeded before the timeout fired, because the retry wrapped a non-idempotent operation with no idempotency key. The test proved the retry ran. It never asked whether retrying was safe to do — and neither did the code, because the agent had no concept that the two are different questions.

That is what "prove the failure case" means in practice. Don't read the guard and nod. Break it and watch the suite. Comment out the branch, flip the boundary from < to <=, delete the check, re-run. If the tests stay green, they were never testing the guard — they were decorating it.

you@project:~ bash (80x24)
you@project:~$ pytest -q tests/test_mailer.py
1 passed in 0.42s
you@project:~$ # revert the retry guard the agent claims fixes the timeout
you@project:~$ git checkout HEAD~1 -- app/mailer.py
you@project:~$ pytest -q tests/test_mailer.py
1 passed in 0.41s   # ← test never touched the guard: decoration

A test that cannot fail is documentation with a green checkmark. Make the agent show you the red line — the same test failing against the code before its fix. No red, no proof.

4. Check what it touched that you didn't ask for. Version control keeps an honest record. A config file, a lockfile, an unrelated module — those unrequested changes are where cascade bugs live.

you@project:~ bash (80x24)
you@project:~$ git diff --name-only
app/handlers.py
tests/test_handlers.py
config/settings.py   # ← didn't ask for this

The payoff

The discipline compounds. Agent code that clears this gate is code you can trust at main — and the next agent prompt is stronger because it learned the bar from the last one. Every guard-check it survives becomes a convention you can write into the repo's contributing notes. You stop reviewing everything and start catching the thing that was going to bite you.

The honest boundary

This is not a substitute for architecture review. A bot can't tell you the API shape is wrong for the long term, or that you're building the wrong thing. What it can do is catch the silent, testable, obvious failures before they reach a customer — which is most of the damage.

Try it once

Take the last agent-authored commit in your repo and run git show --stat HEAD to see its unrequested changes. One command, read-only, and you'll get a feel for how often the summary and the diff disagree.

References

AI-assisted, curated for Side Quest Studios.

💬 Discuss this article in the community — 0 replies →
Comments live on community.sqs.chat — one thread per article.