
The author's 「契約ループ開発」 splits work into three layers — intent, contract, and loop. Humans decide only why something is built and what counts as done; AI handles implementation, verification, and merging.
Summaries like this, in your inbox every morning.
The method is presented as a response to a specific gap: AI writes code faster, but checking whether that code is correct barely speeds up, because verification still falls to people or tests. The author cites three outside signals for this. Andrej Karpathy, in a February 2025 post, called letting AI output flow and forgetting the code exists "vibe coding," and a year later wrote that LLM-agent programming is becoming the professional standard, though oversight and scrutiny increase rather than decrease. The 2025 DORA survey of roughly 5,000 technical workers frames AI mainly as an amplifier, meaning it magnifies both organizational strengths and weaknesses. And "loop engineering," a term the author traces to an arXiv preprint described as under review, designs the system that issues instructions to agents rather than instructing them each time.
The author's reading of those three is flagged as personal interpretation: as AI writes more, the design of the mechanism that verifies and stops the work determines the result. The concrete answer is a fixed set of roles. verify owns pass/fail, stop-gate owns whether the loop stops, merge-gate owns whether code merges. stop-gate runs as a hook when the AI tries to finish — the author uses Claude Code's Stop hook — and if verification fails, it blocks exit and returns the failure to the AI. To keep a stuck loop from spinning forever, stop-gate fingerprints the failure and hands off to a human when the same failure repeats, in this case three times.
Autonomy is set mechanically by the file paths a change touches. Locations tied to authentication or database migration are configured as high risk and always require human approval; changes matching nothing are low. The upgrade criteria are rework rate and post-merge defects over recent merges, counted in an approximate way from commit trailers. The author also describes removal as a first-class practice: every component records which model limitation it assumes, and provisional components are removed one at a time each quarter or when a new model arrives, with evaluation cases rerun. Anthropic's engineering blog is cited as holding a similar view — that components embed assumptions about what a model cannot do alone, and those assumptions go stale.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
IBM's Bruno Aziza said the number of AI agents employees build will outpace what companies can manage, so firm…
Nathan Lambert published an essay arguing AI progress will accelerate through engineering and infrastructure g…

Google DeepMind's Pushmeet Kohli said AlphaFold did not solve protein folding, because proteins are disordered…

ALPHA FORGE's new sandbox.py calls the same inference orchestrator as the daily batch but never calls the ledg…

Writing on Zenn, Taichi Endoh — a clinical engineer and AI engineer — says the first step in a leak is to defi…

The guide walks through creating a TypeScript MCP server with @modelcontextprotocol/sdk, registering a single…
