AIToday
Large Language ModelsAI Safety & AlignmentThe Rundown AIPublished: Sep 16, 2026, 22:00 JST

China researchers' "Last AI" roadmap draws RSI ladder

China researchers' "Last AI" roadmap draws RSI ladder

3 Key Points

  1. What happened

    Thirty-five Chinese researchers, with authors tied to ByteDance, Tsinghua and Shanghai AI Lab, published "The Last AI Built by Humans", a five-level recursive self-improvement roadmap, and classified 491 papers: 75.4% at Levels 1–2, 5.9% at Level 5.

  2. Why it matters

    The authors frame RSI as a milestone, and OpenAI, Anthropic and Google each treat substantial research automation as a risk area, since it could push capabilities beyond effective oversight, according to the paper's account of their frameworks.

  3. What to watch

    The safety picture hinges on a mismatch the body flags: the roadmap measures control over the improvement process, while lab frameworks assess speed, scale and risk, so a Level 5 label alone is incomplete. Anthropic's Responsible Scaling Policy took effect July 8.

WHO IT HITSAI safety and governance teams at labs, plus research-infrastructure and release-approval staff, face harder calls about who checks changes and who can block a release if research speeds up, the body's framing suggests.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The paper's ladder is not a claim that AI has reached self-improvement; it is a scoring scheme. The authors placed 491 existing papers on it, and the distribution they report is bottom-heavy: 75.4% at Levels 1 and 2, where humans still design the upgrades or the AI only proposes fixes, against 5.9% at Level 5, where AI redesigns the improvement process. They also note that sustained gains across successive model generations remain an open problem, so the upper rungs describe a direction rather than a demonstrated capability.

The roadmap arrives alongside lab policies that approach the same territory from a different angle. OpenAI's Preparedness Framework puts fully automated AI research and development at its Critical threshold, with one indicator being sustained improvements across model generations taking one-fifth as long as equivalent work in 2024, and specifies halting further development until safeguards are defined. Anthropic's Responsible Scaling Policy, effective July 8, sets a threshold around replacing its entire research scientist and research engineer workforce at competitive costs, and the company acknowledges the threshold has been hard to apply. Google DeepMind's Frontier Safety Framework sets a threshold at fully automating a Google team focused on improving AI capabilities, at roughly comparable total costs.

Those frameworks measure speed, scale and risk, while the paper's ladder measures control over the improvement process, and the body notes that a Level 5 label alone therefore gives an incomplete picture of the safety question. The comparison also has a national dimension with limits: the paper reflects its authors' outlook, and AP reported that Guo Jiakun, a spokesperson for China's foreign ministry, rejected AI warnings aimed at China while calling for cooperation on global AI governance, suggesting resistance to the geopolitical framing. For labs and safety teams, the practical test is likely to be narrower than any label: how much research a system completes, how long it takes, what it costs and how much human help it needs, and who inside a lab can check changes and block a release.

FAQ
What are the five levels of recursive self-improvement in the paper?
Level 1 has AI carry out upgrades humans designed, and Level 2 lets it diagnose weak spots and decide how to fix them. Levels 3 and 4 hand over decisions about what the model learns next and how it adapts after deployment, and at Level 5 AI can redesign the improvement process itself.
Why do the researchers see coding as the clearest route to RSI?
Software changes can get immediate feedback from tests, while robotics, science and medicine face slower, more costly feedback, according to the paper. The tests still need to establish that a change improved the system.
How do OpenAI, Anthropic and Google treat automated AI research?
OpenAI's Preparedness Framework places fully automated AI research and development at its Critical threshold, one indicator being sustained improvements across model generations taking one-fifth as long as equivalent work in 2024. Anthropic's Responsible Scaling Policy effective July 8 sets a threshold around replacing its research workforce, and Google DeepMind's Frontier Safety Framework sets one at fully automating a Google team focused on improving AI capabilities.
The Rundown AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Salesforce drops UI as a moatStratechery (Ben Thompson) · 2h ago
  • Diogo Almeida's TypeSafe launches Jev: 20–200x faster decision modelLatent Space · 2h ago
  • Salesforce debuts Koa, a homegrown agent model built on Nemotron 3 SuperThe Rundown AI · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleArcee AI hits $1 billion valuation after $20 million open-weight bet