
What happened
AINOW's guide says AI coding agents now run planning, implementation, testing and fixes on their own, citing a Stack Overflow survey where about 31% of developers already use them at work.
Why it matters
The gap is widening between teams that delegate routine implementation and those that wait, the article argues, so design and review time can replace hands-on coding.
What to watch
Gains reported by Rakuten and Nubank are single-company cases, so the test is whether your team can start from small, well-defined tasks and grow that scope safely.
WHO IT HITSEngineering managers and developers evaluating coding tools are the primary audience, since the article's suggested entry point is low-risk work like adding tests or fixing minor bugs. Security, legal and finance teams also have a stake, given the article's warnings on leaked credentials, license-tainted code and usage-based billing.
Summaries like this, in your inbox every morning.
The article frames AI coding agents as a distinct step beyond autocomplete: rather than predicting the next line, an agent keeps a work plan, edits files, runs tests and revises until the task passes. It sorts the tools into three styles — terminal agents like Claude Code that work across a whole repository, IDE agents like Cursor that show diffs for approval, and cloud agents like Devin that take a task and return a pull request. Rule files such as CLAUDE.md and AGENTS.md carry project conventions between sessions, and permission settings decide how much the agent may do without asking.
The evidence the article offers is mostly company case studies. Rakuten had Claude Code work for 7 hours on vLLM, a 12.5-million-line open-source library, and reported cutting lead time from 24 business days to 5. Nubank used Devin on a migration of more than 6 million lines, reporting that work once forecast at 18 months and over 1,000 engineers finished in weeks at more than 20x lower cost. Classmethod reported up to 10x productivity and an 80% cut in code-review time, and ULS Group used Devin on about 1.5 million test steps, cutting effort from 200 person-months to 50. Set against this, Veracode's 2025 survey found 45% of code samples from over 100 AI models failed security tests, and 46% of developers told Stack Overflow they do not trust AI tool output.
That contrast is the crux. The outcome likely hinges on process rather than tool choice — clear task boundaries, limited permissions, rule files and automated test and CI checks. Teams with those in place may be able to widen what they delegate; teams without them risk carrying generated errors into production. The article's own answer to the skeptics is to start small, on tasks where success or failure is easy to judge.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Google is testing "Call for Me," letting Gemini call businesses for US Pixel 11 owners with a paid Gemini subs…

Anthropic signed a seven-year, $11.6 billion cloud deal with Akamai Technologies, per Reuters, and gets a warr…

Microsoft unveiled a redesigned Copilot "super app" combining chat, coding, Office tools, and a new "Autopilot…

Reviewer Jennifer Pattison Tuohy tested Apple Intelligence for Home, Google Nest's Gemini features, and Amazon…

Microsoft officially unveiled its redesigned Copilot super app with three tabs — Home, Code, and Autopilot — a…

Okta announced Monday it is collaborating with CrowdStrike, Google Cloud, Salesforce and others to launch the…
