AIToday
Large Language ModelsAI Safety & AlignmentAhead of AI (Sebastian Raschka)Published: Sep 9, 2026, 22:00 JST1 min read

GPT-6 Astra tops ARC-AGI-3 at 99.9%, looped transformer rumors unconfirmed

GPT-6 Astra tops ARC-AGI-3 at 99.9%, looped transformer rumors unconfirmed

3 Key Points

  1. What happened

    Last week, OpenAI released GPT-6 Astra. The author calls it the best model they have used, with standout performance in 3D rendering and animation tasks.

  2. Why it matters

    Astra scores 99.9% on the ARC-AGI-3 benchmark, versus 7.8% for GPT-5.6 Sol, while also improving across writing, math, and coding.

  3. What to watch

    Whether 'looped transformer' architecture is real hinges on unverified reporting from The Information. The author suggests hiding reasoning traces is a possibility, not a fact.

WHO IT HITSAI researchers and benchmarking professionals will need to weigh independent benchmarks against rumors. Developers of agentic applications may see new computer-use capabilities, but should treat unverified architecture claims with caution.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The article discusses GPT-6 Astra's capabilities, noting its impressive performance in computer use and 3D tasks, which the author attributes to training on macOS environments via thousands of Macs. This aligns with a trend toward making LLMs more accessible for everyday tasks outside tech.

The author then explains looped transformers, citing examples like Nanbeige and Ouro, which reuse blocks to increase depth. Recent research, such as Mixture-of-Recursions, shows this can improve quality at fixed compute budgets for larger models. This context is crucial for evaluating the rumor about Astra's architecture.

The stakes hinge on whether the looped transformer rumor is true. If confirmed, it could explain Astra's efficiency and advanced capabilities. However, with no official confirmation, the community must rely on independent benchmarks to assess the model's true standing.

FAQ
What is GPT-6 Astra's benchmark score on ARC-AGI-3?
GPT-6 Astra achieves 99.9% on the ARC-AGI-3 benchmark, while GPT-5.6 Sol scores only 7.8%.
Why is looped transformer architecture important?
Looped transformers reuse the same blocks to increase depth without adding weights, potentially improving quality at a fixed compute budget if the model is large enough, as seen in examples like Nanbeige4.2-3B.
Ahead of AI (Sebastian Raschka)Read Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DeepSeek V4.1-Flash: 763B model beats V4 Pro on AA Index 40Latent Space · 2h ago
  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 8h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 8h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAI's early impact too indirect for average people, Interconnects warns