AIToday
AI Business & IndustryAI Safety & AlignmentFortune AIPublished: Sep 5, 2026, 10:01 JST1 min read

OpenAI quietly tweaks GPT-6 Astra benchmark scores post-launch

OpenAI quietly tweaks GPT-6 Astra benchmark scores post-launch

Key takeaway

  • OpenAI revised GPT-6 Astra benchmark scores after launch. Some numbers favored Astra, while Anthropic's dropped.

  • OpenAI says fixes ensure accurate comparisons. Critics question the practice of re-running evals.

  • The confusion could affect customer trust and OpenAI's future IPO.

3 Key Points

  1. What happened

    OpenAI changed several evaluation metrics for its GPT-6 Astra model after first publishing a blog post on Sept. 3. The numbers on updated versions showed Astra performing better, while Anthropic's scores got worse.

  2. Why it matters

    The changes highlight concerns about benchmark reliability and possible manipulation, as these scores help AI companies win customers and investors. OpenAI said it made fixes to ensure numbers reflect the best estimate of model performance.

  3. What to watch

    The hallucination rate was halved to 2% for Astra then reverted to 4.2%, while Sol's ExploitBench score rose to 11.5% and may be reverted. OpenAI faces scrutiny over its evaluation practices ahead of a possible 2027 IPO.

Ask the AI about this article →

Context & Analysis

The benchmark adjustments come as OpenAI and rivals race to lead in AI, where eval scores serve as a public scorecard. The changes, including halving Astra's hallucination rate to 2% and boosting Sol's ExploitBench score, could give OpenAI a marketing edge. But experts like Stanford researchers note that re-running evals is a known practice, and the system card provides few details. This raises questions about how much trust to place in these numbers. The confusion could complicate customers' model choices and muddy OpenAI's narrative of having the best models, especially ahead of a potential 2027 IPO.

FAQ

Why did OpenAI change the benchmark numbers?
OpenAI said the changes were fixes to ensure numbers represent the best estimate of available model performance. The company also noted that most evaluations have noise within a few percentage points.
What was the most notable change?
Astra's hallucination rate was halved to 2% in later versions of the blog, then reverted to 4.2%. Also, GPT-5.6 Sol's ExploitBench score rose from 5.5% to 11.5% and may be reverted.

Get the latest AI Business & Industry news every morning

For example, today's edition would include:

  • Topco eyes optical comms and cooling as copper fadesDIGITIMES Asia · 54m ago
  • Microsoft to Rework Reporting for AI, Cloud From FY2027Yahoo Finance AI · 54m ago
  • ChronoScale Plans 50 MW Microsoft AI BuildYahoo Finance AI · 54m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGPT-6 Astra Early Users Report Overly Cautious Tone