AIToday
Large Language ModelsAI Coding AssistantsZenn AI/MLPublished: Oct 1, 2026, 22:00 JST

Adversarial AI code review: useful on 99 回, but 9× the tokens

Adversarial AI code review: useful on 99 回, but 9× the tokens

3 Key Points

  1. What happened

    A team ran its own Claude Code adversarial-review skill for 99 reviews of 85 diffs, using 8 parallel lanes plus a YAGNI lane, and confirmed 357 findings — 138 Critical or Important, 110 Suggestion and 97 Nit.

  2. Why it matters

    The extra cost looks steep next to Claude Code's standard code-review, which caught only 7 of 100 confirmed Critical or Important findings, while adversarial review found 72 Critical issues via a single lane only.

  3. What to watch

    The test was one team's own workflow, and 16 of 99 runs confirmed nothing, all on diffs under 500 added lines — so small, non-critical changes may not justify the extra tokens.

WHO IT HITSEngineering teams and developers adopting AI code review must weigh the 9× token cost against the extra Critical findings, especially on diffs touching data updates, calculation logic, or authorization.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Adversarial review works by splitting a diff among several agents with different roles — a conventional reviewer, a bug hunter, and lanes that deliberately hunt for inputs or states that break the change — then cross-checking their claims against the actual code before confirming anything. The team built its version as a Claude Code skill with 8 lanes across three families, plus a separate YAGNI lane that drops or downgrades findings whose real-world harm cannot be shown. Every review logged which lane raised how many findings and how many stuck.

The numbers show both sides of the trade-off. Cross-validation passed 682 findings, and the YAGNI lane cut them to 357 confirmed, so a large share of the raw output never reached a human. The heaviest cost comes before that filter: the main agent must cross-check 1,413 lane findings to reject 1,056 of them, and even discarded findings consume time and tokens. The team also added rules to keep the loop from running away — separating who raised a finding, capping automatic fix cycles, and letting Fable settle split judgments.

Read against the author's conclusion, the practical value seems to depend less on diff size than on what the change touches. A diff under 100 added lines still produced Critical findings in 9 of 23 runs, while nothing was confirmed on any diff under 500 lines in the zero-finding runs — so the deciding factor may be whether the change touches data updates, calculation logic, or authorization. Teams weighing adoption may find the payoff hinges on routing only those critical paths through adversarial review, though the token overhead of cross-validation is likely to remain the main constraint.

FAQ
How does adversarial review compare with Claude Code's standard code-review?
In the team's comparison, standard code-review detected only 7 of the 100 confirmed Critical or Important findings when both ran, though it alone added 7 findings across 82 runs. The author notes code-review focuses on changed lines, so it structurally misses issues that require comparing against prior code or other files.
When did adversarial review confirm zero findings?
In 16 of 99 reviews it confirmed nothing, and all of those diffs added fewer than 500 lines. The author says that at that size, adversarial review often yields nothing.
How were disputed findings resolved?
A separate Claude model, Claude Fable 5, judged the disagreements. Of 99 reviews, 31 had Fable requests covering 64 items: 26 judged problematic, 27 not, 7 downgraded, and 4 other.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOllaya runs decision models locally, up to 255 options