AIToday
Large Language ModelsAI Coding AssistantsZenn AI/MLPublished: Oct 6, 2026, 22:00 JST

Codex reviews Claude Code script: 4 issues, then 1, then 0

Codex reviews Claude Code script: 4 issues, then 1, then 0

3 Key Points

  1. What happened

    A first-year engineer had Claude Code write a small script checking pages for a noindex tag, then had Codex review the uncommitted change three times: 3 issues plus a supplement round one, 1 new issue round two, 0 issues round three.

  2. Why it matters

    The review paid off because the author had only tested two prepared sample cases, so the flaws that slipped through were in the cases the author had not tried; fixing one issue also opened a fresh hole.

  3. What to watch

    The whole setup hinges on the reviewer staying hands-off — the engineer used a read-only sandbox and compared git status before and after to confirm nothing changed. Codex explicitly flagged Japanese filenames and non-UTF-8 HTML as unverified in all three rounds.

WHO IT HITSDevelopers and small engineering teams who run two AI coding assistants side by side and want a cheap second pair of eyes on code before committing it.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The engineer's own write-up is notable for showing how the two tools were kept from sharing context: Claude Code and Codex do not share conversations, so the handoff note at docs/handoff.md carries only three items — current state, what changed, and the next request. The note limited the reviewer's scope to uncommitted changes and told it not to fix anything, leaving the decision on which points to accept with the author and Claude Code. Codex did not just return opinions; it wrote small tests to reproduce each point before reporting, without touching the file.

Across the three rounds the reviewers' notes never appeared in the author's own check, which had covered only two prepared sample cases. That said, the engineer is explicit that this is a single small script and cannot be generalized. The test of whether the approach is worth copying is whether review catches what the author did not try, and whether the author can keep the reviewer read-only.

FAQ
How did Codex avoid changing the files it reviewed?
The engineer ran Codex in a read-only sandbox with the -s read-only flag, so commands it generated could not write, and compared git status before and after.
How many times was the script reviewed, and how many issues came out?
Codex reviewed it three times. Round one yielded 3 issues plus 1 supplement, round two yielded 1 new issue, and round three yielded 0 issues.
What did Codex flag as not yet verified?
Codex stated in all three rounds that Japanese filenames and HTML that is not UTF-8 had not been verified.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGoogle locks Gemini Pro behind AI Pro as free tier drops to Flash-Lite