
What happened
The engineer wired Claude Code headless into a pipeline that turns backlog items into merged code, logging 243 tasks (239 done) and 325 pull requests, all merged. His sales division of 25 AI roles launched September 16, 2026.
Why it matters
The setup shows AI can carry routine dev and sales work, but his worst failure — an AI reporting 'already implemented' while the harness deleted the finished branch — suggests self-reported success is unreliable.
What to watch
His fixes hinge on verifying commits and tests instead of AI claims; the open question is how far autonomy spreads, since only low-competition, high-scoring deals are auto-selected and sending still needs human approval.
WHO IT HITSThis lands on engineering leads and solo developers considering AI-driven development pipelines, and on operations teams automating outbound work — the record shows where human approval gates and outcome checks must sit.
Summaries like this, in your inbox every morning.
The write-up is unusual less for its tooling than for its honesty: it is a mid-flight repair log, not a finished product pitch. The engineer says he built the development harness in July 2026 and the sales division on September 16, 2026, moving the latter from a Mac to an always-on small PC on October 5, 2026, while dropping a contact-form outreach experiment the same day after watching the response numbers. The two systems share one design choice — a cloud ledger that both the admin screen and each AI role read and write — which let the machine swap happen without touching the screen, though running two machines at once would double-send.
The most instructive thread is what he calls the pitfalls, and each follows the same arc: a symptom, a cause, a fix. A task was marked complete as "already implemented" because success was judged by whether the working tree held uncommitted changes, so an agent that committed on its own looked empty; the harness retried, accepted the AI's own report, skipped review and merge, and cleaned up the very branch holding the work — recovered only because the commits still existed unreferenced. Budget caps produced a second trap: when a limit hit before the final exchange, a run could end with a clean exit code yet no output, so he now classifies both cutoff shapes as over-budget and preserves partial progress as temporary commits instead of restarting from zero.
The database episode shows the stakes of running on a free tier — the admin screen's full-page reload every 30 seconds, indexed to read only a small change marker, now costs 114 rows. Whether this approach generalizes remains open: the harness has so far mostly developed itself, and its wider claims rest on a handful of personal apps, with the human-judgment rules — approval queues, morning triage, the two-strike limit on legal rewrites — still the part doing the load-bearing work.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Cline Bot CEO Renee Huang said the company's open-source coding harness is expanding beyond coding into genera…
OpenAI released its Jev-style Decisions API, previously announced at last week's DevDay, and Simon Willison us…

On September 29, 2026, Anthropic published a cyber-capability and safety evaluation of Z.ai's GLM-5.3, reporti…

OpenAI says it will automatically watermark ChatGPT text in the European Union and offer the feature elsewhere…

Lucas Baker, Head of LLM R&D at Jump Trading, said GPT-6 Astra lets agents run multi-day quant research workfl…

Fujitsu announced four AI agents for retail store operations and management decisions, covering sales structur…
