Review Habits for AI-Generated Code

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.
Review an agent-written change by checking the requested behavior, the final diff and the observed verification results. A short handoff can make those artifacts easier to find. It cannot establish correctness by itself, even when the agent describes the work confidently.
Updated 22 September 2026: combined three overlapping review checklists, removed the unrelated product-blog framing, and clarified the difference between instructions and enforced permissions.
This is our proposed review exercise for a small repository change. It applies whether the author used Cursor, Claude Code, Codex or another assistant. It is not a claim that a particular tool has been evaluated or that every change needs the same amount of process.
Start with a failure a reviewer can observe
Choose an existing parser bug, such as accepting an option without its required value. Write down the expected error and one ordinary valid-input case before asking the agent to repair it. Keep the task small enough that another person can inspect every changed file.
After the agent finishes, compare the implementation and tests with that original requirement. A test that merely repeats the new implementation's output can pass while preserving the bug. Look for a meaningful assertion about the expected behavior, and check whether existing assertions were removed or weakened.
Inspect the full working tree, including new files. An unexpected dependency, generated artifact or configuration change can matter even when the main function looks correct. Split unrelated cleanup from the behavior change when doing so makes review easier.
Check that verification belongs to the final patch
Require the exact command and its outcome. Check whether the last successful run happened after the last relevant edit. If a command failed, retain that fact and explain whether a later run resolved it. A statement that tests should pass is a prediction, not an observed result.
Use the repository's normal test and integration paths. For changes involving persistent data, authorization or external effects, include the relevant failure behavior and recovery plan. A small local unit test cannot establish that a deployment or migration is reversible.
Review screenshots or manual observations when the behavior is visual, while keeping the limits clear. A screenshot captures one state; it does not cover every interaction. Name what was checked and what remains uncertain.
Make the handoff short and verifiable
| Field | Reviewer question |
|---|---|
| Task | What behavior was requested? |
| Scope | Which files changed, including unexpected additions? |
| Evidence | Which command or observation checks the final result? |
| External context | Which source, issue or tool result affected the patch? |
| External actions | Were any remote records or systems changed? |
| Uncertainty | Which assumption or skipped check still matters? |
| Decision | Who accepts, revises or rejects the change? |
Keep recurring test commands and conventions in the repository's instruction files. Treat them as guidance to the agent, not an access-control system. Restrict credentials, filesystem access and write-capable integrations through the components that actually execute actions. A note saying read-only does not disable a write tool.
Use this handoff on a few real changes and remove fields that add no review value. Add a sharper check when a concrete failure exposes a gap. The goal is a patch another maintainer can understand and verify without reconstructing an entire agent conversation.