
What happened
Anthropic tested Claude Code handling daily software maintenance on its own in-house apps for the past few weeks. Claude created 388 pull requests across Anthropic's repositories; 180 were merged after human and automated review, a merge rate of about 46 percent.
Why it matters
The test shows that AI can reduce the burden of repetitive maintenance work on human developers. Claude runs 12 specialized routines—covering crash fixing, dead-code removal, duplicate merging, test cleanup, and logic simplification—using plain-language Slack prompts with no elaborate prompt engineering required.
What to watch
More than half the auto-generated pull requests did not merge, signaling both limits and potential of the approach. Anthropic engineer Boris Cherny describes the results as "surprisingly positive" and "early signs of life" for autonomous AI-powered maintenance, and the company is exploring ways to speed up the merge process for mechanical changes.
Ask the AI about this article →
Anthropic's experiment with Claude Code represents a pragmatic exploration of how AI can absorb repetitive software maintenance tasks. Rather than a theoretical test, the company deployed Claude against its own codebase, running 12 distinct maintenance routines—from crash fuzzing to duplicate-code merging to flaky-test fixing—each with a specific role in the broader maintenance pipeline. The setup is deliberately unglamorous: plain-language Slack prompts, no complex prompt engineering, and straightforward human review gates. This approach avoids over-engineering the system while keeping humans in control of what ships.
The 46 percent merge rate after review is telling. It indicates that Claude can produce valid, useful pull requests at scale—enough to meaningfully reduce developer toil—but also that nearly half the proposals need refinement or rejection. The team's iterative approach (tweaking routines when they fail) suggests a path toward improvement over time. Cherny's framing as "early signs of life" rather than "solved" is honest and important: the experiment shows potential but does not yet prove that fully autonomous maintenance is ready for production. The fact that Anthropic is already looking at ways to accelerate the merge pipeline suggests confidence in the direction.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Analyst Ming-Chi Kuo says Nvidia has revived the Rubin CPX AI accelerator with a substantially redesigned arch…

A UK study by UK AI Security Institute and Limbic AI surveyed 6,474 British adults

Broadcom's Clayton Donley says companies are doing mission-critical work with AI agents quickly, but without t…
OpenAI released a new evaluation framework on July 17, 2026, urging companies to measure AI ROI by 'useful out…

As AI agents perform real business tasks, 'Agentic Identity' (giving each AI a unique employee-like ID) and 'D…

The European Union is expanding regulation of ChatGPT and will mandate protections for minors
