When AI Joins the Pull Request, Review Depth Still Comes From Risk

One wiki pull request received 13 findings in its first hosted review round. The fixes landed, and the local review drain became clean.

The next hosted round found seven more defects. All seven had been introduced by the first fixes, and six were contradictions with nearby sentences on the same pages.

Later rounds found issues in source memories written during the pull request and an octopus-merge defect in a hook change. The hook defect was caught just before merge and corrected afterward in a follow-up pull request.

The important variable was not whether a person or an AI wrote each line. The change surface kept moving while review fixes added new claims and new code.

That incident gave me a simpler rule: review depth comes from the change surface and the consequence of being wrong. The author can be a person, an agent, or a particularly ambitious shell script (the script still does not get feelings).

A review mechanism needs a routing rule

Without an explicit threshold, “get another review” expands indefinitely. More reviewers feel safer, even when they read the same files with the same assumptions.

The operating policy I use starts with three physical tiers:

  • Small: at most 5 files and 200 changed lines, with no core area involved.
  • Medium: 6 to 15 files or 201 to 600 changed lines.
  • Large: more than 15 files or more than 600 changed lines.

A core-area change moves up one tier. Security, authentication, migrations, concurrency, release automation, and similar surfaces deserve deeper scrutiny even when the diff is short.

These numbers are policy thresholds, not a scientific optimum. They make the routing decision visible and adjustable. The source record does not measure whether this exact split detects more defects per token.

That honesty matters. A useful heuristic does not need to pretend it is a benchmark.

Small changes need context, not ceremony

A small, non-core change gets an inline review.

The reviewer reads the request, inspects the full diff, checks relevant call sites, and examines the focused verification. Spawning several independent reviewers for a two-line copy fix usually adds coordination noise rather than new evidence.

Small does not mean automatic approval. A one-line permission change can be core and move up immediately. Size is the starting signal, not the final verdict.

Medium changes need one structured pass

A medium change benefits from a consistent checklist:

  1. Does the diff implement the requested behavior?
  2. Which public or internal contracts changed?
  3. Do tests exercise the failure state as well as the success state?
  4. What platform, dependency, or runtime paths remain unverified?
  5. Does the summary match the actual diff?

One reviewer can answer those questions if the scope remains coherent. Adding more agents before the first structured pass often duplicates file reading.

Large changes benefit from independent boundaries

Changes that reach the Large tier after any core-area promotion get parallel review plus an adversarial cross-check in this policy.

Parallel does not mean “ask three models whether the code looks good.” Each reviewer needs a separate evidence boundary. One can inspect behavior and tests. Another can inspect security or migration risk. A third can challenge completion claims and search for untested states.

The adversarial pass then tries to falsify the proposed conclusion:

  • Find a requirement the implementation silently dropped.
  • Find a passing test whose fixture cannot produce the failure.
  • Find a gate that does not read the changed platform.
  • Find a claimed deployment that is only a merged pull request.
  • Find a compatibility statement supported by one local toolchain.

This structure increases independence. Model variety can help, but changing the model name without changing the question does not create a new review boundary.

AI authorship does not remove human ownership

GitHub's current Copilot agent workflow ends with a familiar instruction: review the code changes yourself as you would for any contributor's pull request.

That is the correct ownership model.

An agent can open the pull request, run checks, respond to comments, and even review another agent's change. The repository owner still decides whether the evidence is sufficient for the consequence.

The same rule applies to a human-authored pull request. Familiar authorship should not turn a large migration into a shallow review.

Review the evidence claim, not only the code

Agent-generated summaries are often clean enough to hide uncertainty. I treat each completion claim as a review target.

“Tests pass” should identify the command and what it covered. “Works on iOS” should distinguish static inspection, simulator execution, and a physical-device result. “Deployed” should identify the public artifact or endpoint, not only a merged branch.

This is where adversarial review earns its cost. It asks whether the evidence can support the sentence, not whether the sentence sounds reasonable.

What PR #34 changed in my review process

Before assigning reviewers, I calculate the scope and name the dangerous surfaces.

Then I choose the smallest review mechanism that can independently inspect them.

I also re-read the whole edited page after applying a review fix. PR #34 showed why diff-hunk inspection was insufficient: a corrected sentence could disagree with a sentence a few lines away while every local gate remained green.

An extra review round is not automatically valuable. In this case, independent whole-page passes found newly introduced contradictions and a merge-hook defect that the preceding local drain did not see.

That is the result I want the tiering policy to buy: more independent reading when the change is large enough to keep changing under its own fixes.