AIToday
Large Language ModelsAI Safety & AlignmentITmedia AI+Published: Sep 7, 2026, 13:00 JST1 min read

AI benchmark v4.2 adds private tests to curb gaming

AI benchmark v4.2 adds private tests to curb gaming

3 Key Points

  1. What happened

    Artificial Analysis updated its Intelligence Index to v4.2 on September 4, adding two new tests: AA-Briefcase for knowledge work and GDP.pdf for PDF reading. It also removed GPQA Diamond because scores were saturated.

  2. Why it matters

    Private test data now makes up 40% of the index, double the share in v4.1, to prevent AI developers from optimizing for benchmark questions. A further increase is planned for version v5.

  3. What to watch

    The update's effectiveness hinges on whether the higher private-data share truly deters benchmark gaming. Watch for the v5 release, where the share is expected to rise further.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Artificial Analysis's update reflects a broader challenge in AI benchmarking: as models improve, public test sets become less useful because developers can optimize for them. By increasing the share of private data to 40%, the company aims to keep its index meaningful for comparing frontier models like Claude Fable 5.1 and GPT-6 Astra, which now top the ranking. The addition of more realistic tasks, such as knowledge work simulations and PDF reading, also aligns with the goal of measuring practical capabilities rather than test-specific tricks.

The decision to exclude GPQA Diamond, where scores have plateaued, is a pragmatic response to a test that no longer discriminates between top models. The planned further increase in private data share for v5 suggests that this is an ongoing effort. The index's credibility depends on whether the private data actually prevents gaming, which is hard to verify from the outside. For businesses using benchmarks to choose AI vendors, this update is a step toward more trustworthy comparisons, but it remains to be seen if it fully solves the problem.

FAQ
What new tests were added in the v4.2 update?
Two tests were added: AA-Briefcase, developed by Artificial Analysis to measure knowledge work, and GDP.pdf, developed by Surge AI to measure PDF reading ability.
Why was GPQA Diamond removed from the benchmark?
It was removed because model scores on that test were saturated, meaning they had reached a level where the test no longer differentiated between models.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Fukushima Prefecture expands generative AI to all 6,000 staff despite low usageITmedia AI+ · 1h ago
  • Azoma: Brands Must Optimize for 7 AI Shopping AgentsYahoo Finance AI · 1h ago
  • OpenAI, WAN-IFRA, AIRPPU launch AI program for Ukrainian newsroomsOpenAI Blog · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNvidia to Buy Hugging Face for $12.9B