
What happened
From 2026-10-06 the author ran 21 tasks through a loop where dots posts request files, a watcher launches Claude Code headless every 10 minutes, and Codex CLI audits the output. 16 ran unattended; Codex marked 8 GO and 9 needs-fixing.
Why it matters
Tasks written by Claude Code and declared "complete" by it still showed holes when a different model reviewed them, so the audit step appears to be catching a real gap rather than rubber-stamping.
What to watch
The author ran this for only about a day and says the sample is too small to know whether that near-half rejection rate holds, and warns against adopting it broadly. Official dots APIs or webhooks would prompt a redesign.
WHO IT HITSThis lands on developers and solo operators who already run AI coding agents and are weighing whether to add an automated handoff between them; it suggests the value sits in the cross-model audit step, not the automation alone.
Summaries like this, in your inbox every morning.
The setup is deliberately modest in scope. The author describes a loop in which dots decides the policy and writes requests, a watcher script (hub_runner.py) picks them up every 10 minutes and launches Claude Code unattended, Codex CLI audits the finished result, and the outcome is reported back into the dots conversation. The stated division of labor is that dots handles operations while a human keeps the boundaries, with parallelism, model choice, audit on/off, timeout, ordering, and pausing all set by dots through a policy file.
What the author emphasizes most is the friction that only shows up in a scheduled environment. dots cannot write outside its own working folder, so a fixed entry point receives nothing and the entry must be a glob. On Windows, a CLI installed through npm inside the Claude desktop app is isolated under a package path, so it works by hand but fails every time from Task Scheduler unless launched by full path. Two Windows-specific traps are noted as well: os.kill(pid, 0) does not work as a liveness check, and a scheduler may kill the child process along with the parent. The author's completion bar is correspondingly strict, requiring that a request placed by dots itself be picked up through the scheduler and audited, not just run manually.
The stakes here come down to whether the audit loop is genuinely doing work or just adding a step. The measured rejection count suggests a separate model is catching gaps the author's own agent missed, but one day and 21 tasks is a thin basis, which the author acknowledges. The outcome likely hinges on whether that pattern holds over more runs, and whether an official dots API or webhook arrives that would let the file-and-screen-scraping approach be replaced.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Kore.ai Inc. launched Autoloop, which adjusts AI agents toward goals customers set for task completion, busine…
OpenAI launched a Decisions API that evaluates text, images, or both, returning yes/no probabilities, category…

Google launched Playground, a browser-based AI platform that lets adults in the US create games using only tex…

Google made its SynthID detector available globally to anyone with a Google, OpenAI, or Apple account, with a…

In August, DoorDash emailed Bay Area restaurants warning they may be listed on Bites without consent, with ter…

At MIT Future Fest, Tony Fadell said the Rabbit R1, Humane Ai pin, and Limitless pendant failed because they "…
