Getting a Claude Code review worth reading

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.
A Claude Code review can run on your working tree before you push, or on a pull request after. Either way, the default behaviour is to be helpful, and "helpful" for a language model means finding something to say about every hunk of the diff. That is how you end up with a review that mixes a real race condition with a suggestion to rename data to payload.
Tell it what a comment is allowed to be about. In practice a short instruction changes the output more than any model setting:
Review this diff for correctness, data loss, and security only. Do not comment on naming, formatting, or structure.
Put the same rule in your project instructions file so it applies without being retyped. Compare the resulting findings with a review without that instruction on the same diff. Count verified issues and false positives; do not assume fewer comments means better coverage.
What model review is genuinely better at
Not judgement. Stamina. Both people and models can miss relevant context in a long diff. Automated review can provide another pass, but it does not guarantee uniform attention or complete coverage. So the wins cluster in the places human attention runs out:
- Test files, where reviewers nod along. Assertions that check a mock rather than behaviour show up constantly.
- Error branches that were written and never wired to anything.
- A signature changed in one place and left stale at the two call sites nobody grepped for.
- Interpolated values landing in SQL, a shell command, or a URL.
- Boundary conditions in loops and slicing.
Where a Claude Code review falls down
A diff-only review may miss whether the change should exist at all. If the approach is wrong, you get a well-reasoned polish of a bad idea. That failure is invisible if you treat the review as a gate, because nothing red appears.
Without repository history and decision records, it may not know that this module caused last quarter's incident, that the config flag being added duplicates one added under a different name in March, or that your team decided against this pattern for reasons never written down. Some of that you can fix by writing the decisions down. Most teams do not.
It is also not deterministic. Two runs on the same diff give overlapping but different comment sets. If you plan to block merges on it, understand that you are blocking on something that can change its mind. We recommend advisory, not blocking, until you have a month of data.
And it degrades on large changes. A forty-file refactor produces vague findings. That is information about your PR size as much as about the tool.
Fit it around the human review, not into it
The sequence that works is: the agent reviews before a person ever looks. The author runs it against their own diff, fixes what is real, and opens the PR clean. Then the human reviewer spends their attention on design, on whether the change belongs, on what it will do to the on-call rotation. The reviewer remains responsible for those decisions, even when the model contributes useful analysis.
If automated review runs alongside human reviewers, assign someone to triage its findings. Otherwise a comment can remain unresolved because each reviewer assumes another person checked it.