AIToday
Large Language ModelsAI Coding AssistantsZenn AI/MLPublished: Oct 9, 2026, 10:00 JST

Claude Code harness: 203 of 239 AI tasks passed first review

Claude Code harness: 203 of 239 AI tasks passed first review

3 Key Points

  1. What happened

    The engineer ran Claude Code headless, one process per stage, isolating each task in its own git worktree; of 239 completed tasks, 203 cleared the first review, 30 the second and 4 the third.

  2. Why it matters

    Splitting stages and limiting each one's permissions appears to keep a solo developer's codebase workable while AI handles implementation, with humans pulled in only when retries run out.

WHO IT HITSSolo developers and small engineering teams experimenting with AI coding agents are the clearest audience, since the design shows one person can supervise a queue of AI-run tasks with human handoffs only on repeated failure.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The harness rests on a small set of rules the engineer states at the end: never hand everything to one run, split the stages and keep each stage's permissions minimal; use the same session for fixes that need context but a different session and model for independent judgment; force AI verdicts into a fixed shape rather than parsing prose; classify cutoffs by reason and don't let failures that aren't the AI's fault burn through retries; and keep state outside the process so work can resume wherever it stopped.

Several design choices follow from specific incidents. The AI is not allowed to push, merge or switch branches because it had failed several times pushing while the remote was ahead. Design runs read-only and review is limited to Read, Grep and Glob, on the reasoning that a reviewer who fixes things itself leaves no record of what it flagged. Design is deliberately not a mandatory gate: an empty design output does not fail the task, so a stalled auxiliary stage can't halt the main work.

Recovery is treated as a first-class concern. Resume points are recorded after implementation, after review approval and just before merge, so an interruption doesn't restart from scratch. Each launch also reconciles unfinished tasks that have pull requests against their real GitHub state, marking already-merged ones complete so the same task isn't implemented twice. Of completed tasks, 217 finished on the first run and 22 completed only after re-execution.

FAQ
How does the harness stop AI from breaking a developer's local work?
Every task gets its own branch and git worktree created from the development branch, so the AI never touches the developer's working tree or uncommitted changes. The AI is also barred from pushing, merging or switching branches.
Why is the review done by a different model?
By default a higher-tier model than the one used for implementation handles review, because the engineer judged that the same model tends to approve its own output too leniently. Only blocking issues are sent back, and empty or empty-bodied responses are retried rather than treated as verdicts.
Why is GitHub Actions not used for CI?
The engineer wanted to avoid a mechanism that costs money. Test and review results are instead recorded on the pull request via GitHub's commit status API.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleHex cuts wrong edits 7x with evals on Quick Edits