Running a Claude Code security review that finds real bugs

By Rogier Muller08.15.26
Running a Claude Code security review that finds real bugs

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.

It is a model reading your code with a security-shaped prompt. Run /security-review in a session and it examines your pending changes and reports findings. There is also a GitHub Action form of the same idea that comments on pull requests. Underneath, both are the same thing: read the diff, trace where untrusted input goes, report.

Understanding that it is prompt plus diff, and not a static analyser, explains everything about how it behaves. It has no rule database. It is not deterministic. Run it twice and you get overlapping but different findings. And it can reason about code that a rules engine cannot parse, which is the actual advantage.

Where it is genuinely strong

  • Taint tracking through your own helpers. A request parameter that reaches a query builder five functions later, through a wrapper that a pattern matcher would not follow.
  • Authorisation gaps. A new endpoint that checks authentication and forgets to check that the record belongs to the caller. Test for this explicitly with two users or tenants.
  • Secrets and tokens ending up in logs, error messages, or exception payloads sent to a third party.
  • Unsafe deserialisation and template rendering with user-controlled input.
  • Missing output encoding in one branch of a component where the other three branches do it correctly.

The last two categories matter because they are the ones humans skim past. A reviewer reads the interesting part of a diff carefully and the repetitive part quickly. A model can also miss a branch or lose relevant context; another pass is not a completeness guarantee.

Where it fails, and how

It reviews what changed. A diff-only review may miss whether the feature should exist, whether the data should be collected, or whether a tenancy flaw predates the change. Design review is still a person's job.

False positives cluster in predictable places. Test fixtures with hardcoded credentials get flagged as leaked secrets. Internal admin tooling gets flagged for missing rate limits. Code that is protected by a gateway or middleware the model cannot see gets flagged for missing auth. That last one is the expensive one, because the fix is to tell the model where the protection lives, not to add a second check.

It also has no severity calibration of its own. Everything arrives sounding urgent. Somebody has to triage, and if nobody does, engineers learn to close the report unread.

Making the findings worth reading

Give it the context that changes the answer. In CLAUDE.md, write down where authentication and authorisation are enforced, which directories are test-only, what your trust boundaries are, and which sinks are already safe by construction. For example, if these statements match the actual implementation:

Auth: all /api routes pass through requireSession in middleware.ts. Tenant scoping is enforced in db/scope.ts, never in handlers. Anything under fixtures/ is test data, not production secrets.

Then narrow the ask. A review of a whole repository produces a long, shallow list. A review of one pull request produces findings you can act on today. Run it on the diff, before a human reviews, so the human never spends attention on a missing encode.

For a first trial, review one historical pull request with a known authorization defect and one without it. Check whether the report identifies the real access path, then reproduce each finding against a disposable fixture. Track verified findings, false positives and missed known defects. A low useful-finding rate can reflect model limits, missing context or the review task; it does not identify one cause by itself.