AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Sep 6, 2026, 04:00 JST2 min read

AI Index overhauled after GPT-6 Astra scoring skepticism

AI Index overhauled after GPT-6 Astra scoring skepticism

Key takeaway

  • Artificial Analysis updated its AI ranking after criticism that GPT-6 Astra scored too low.

  • The new index gives Astra a four-point gain, placing it second.

  • Claude Fable 5.1 still leads, with Meta third.

3 Key Points

  1. What happened

    Artificial Analysis released version 4.2 of its Intelligence Index, likely in response to criticism that its benchmarks failed to capture GPT-6 Astra's actual progress. With the update, GPT-6 Astra now shows a four-point gain over its predecessor, Sol.

  2. Why it matters

    The revised index places Claude Fable 5.1 first, GPT-6 Astra second, and Meta third. It also adds two new benchmarks—AA-Briefcase for real-world knowledge work and GDP.pdf for PDF analysis—and makes private test data 40 percent of the weighting to make gaming harder.

  3. What to watch

    The test is whether the lead ranking holds, since the two OpenAI models were previously tied and the gap is only four points. Watch the 40 percent weighting of private test data that now underpins the revised scores.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The revision comes after earlier evaluations, including OpenAI's own, showed Astra far ahead of the field, while Artificial Analysis initially scored it just on par with its predecessor. This discrepancy drew skepticism, prompting the index overhaul. The new version adds benchmarks like AA-Briefcase for real-world knowledge work and GDP.pdf for PDF document analysis, and drops GPQA-Diamond because models have solved it. By making private test data 40 percent of the weighting, the index aims to resist gaming.

For business readers, the update matters because it provides a more accurate picture of frontier model capabilities, which is useful for procurement and strategy decisions. The fact that Astra now shows a four-point gain and uses fewer tokens per task suggests cost and efficiency advantages, which could influence vendor choices. However, as the index is still evolving, and version 5 will roll out in stages, these rankings may continue to shift. The top of the leaderboard is moving quickly, and this was an interim adjustment to keep pace.

FAQ

Why did Artificial Analysis change its index?
The update follows criticism that the previous benchmarks did not capture GPT-6 Astra's real progress. The company also fixed scoring errors and added new benchmarks.
What is the new ranking of top models?
Claude Fable 5.1 leads, GPT-6 Astra is second, and Meta is third. Both OpenAI models were previously tied.
How much more efficient is GPT-6 Astra?
According to Artificial Analysis, Astra uses fewer tokens per task than every other frontier model.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • OpenAI reveals AI agents accelerating research at 3.1× human paceITmedia AI+ · 28m ago
  • OpenAI agents hack German site, incident undisclosedSemafor Tech · 28m ago
  • US-China AI gap narrows as costs divergeNikkei AI Stocks · 28m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleOpenAI admits wiki incident, unveils new reporting framework