
What happened
A team ran its own Claude Code adversarial-review skill for 99 reviews of 85 diffs, using 8 parallel lanes plus a YAGNI lane, and confirmed 357 findings — 138 Critical or Important, 110 Suggestion and 97 Nit.
Why it matters
The extra cost looks steep next to Claude Code's standard code-review, which caught only 7 of 100 confirmed Critical or Important findings, while adversarial review found 72 Critical issues via a single lane only.
What to watch
The test was one team's own workflow, and 16 of 99 runs confirmed nothing, all on diffs under 500 added lines — so small, non-critical changes may not justify the extra tokens.
WHO IT HITSEngineering teams and developers adopting AI code review must weigh the 9× token cost against the extra Critical findings, especially on diffs touching data updates, calculation logic, or authorization.
Summaries like this, in your inbox every morning.
Adversarial review works by splitting a diff among several agents with different roles — a conventional reviewer, a bug hunter, and lanes that deliberately hunt for inputs or states that break the change — then cross-checking their claims against the actual code before confirming anything. The team built its version as a Claude Code skill with 8 lanes across three families, plus a separate YAGNI lane that drops or downgrades findings whose real-world harm cannot be shown. Every review logged which lane raised how many findings and how many stuck.
The numbers show both sides of the trade-off. Cross-validation passed 682 findings, and the YAGNI lane cut them to 357 confirmed, so a large share of the raw output never reached a human. The heaviest cost comes before that filter: the main agent must cross-check 1,413 lane findings to reject 1,056 of them, and even discarded findings consume time and tokens. The team also added rules to keep the loop from running away — separating who raised a finding, capping automatic fix cycles, and letting Fable settle split judgments.
Read against the author's conclusion, the practical value seems to depend less on diff size than on what the change touches. A diff under 100 added lines still produced Critical findings in 9 of 23 runs, while nothing was confirmed on any diff under 500 lines in the zero-finding runs — so the deciding factor may be whether the change touches data updates, calculation logic, or authorization. Teams weighing adoption may find the payoff hinges on routing only those critical paths through adversarial review, though the token overhead of cross-validation is likely to remain the main constraint.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
ServiceNow launched Flow, a natural-language service desk that handles employee requests inside Slack and Micr…
DeepSeek and Huawei announced a partnership to develop semiconductor software, part of efforts to reduce China…

The AI Ataraxos beat Niemeijer, the most decorated Stratego player, with an 85 percent effective win rate over…

Google unveiled Gemini 4 Argon on September 30, saying it beats GPT-6 Astra and Claude Opus 5.5 on 13 of 19 be…

Yann LeCun told Fortune’s Emily Forlini he has zero concerns about rogue AI incidents, including OpenAI agents…

A Preply survey of over 5,000 professionals in nine countries found 92% of Gen Z respondents used AI for learn…
