AIToday
Large Language ModelsAI Safety & AlignmentZenn AI/MLPublished: Oct 11, 2026, 10:00 JST

Paper: 56 AI Benchmarks Found Widely Mismatched

Paper: 56 AI Benchmarks Found Widely Mismatched

The 2026 paper "What AI Benchmarks Actually Measure" examined 56 benchmarks — 17 capability-related and 39 safety-related — across 53 models, and found safety benchmarks under the same label often ranked models differently, while different capability labels such as reasoning, knowledge and reading comprehension often produced similar rankings.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The paper's approach borrows from educational and psychological measurement, where "validity" asks whether a test measures what it claims to. Stanford University's teaching materials on AI measurement science frame AI evaluation similarly, as a measurement problem that estimates latent abilities from observed responses rather than simple score-keeping, and NIST has published practical guidance treating evaluation conditions, analysis and reporting as one process.

The findings have implications beyond AI. Human university entrance exams also split multifaceted ability into subjects, and combine different test types into a single pass/fail judgment; the Ministry of Education calls for multifaceted evaluation while demanding fairness and clarity of criteria. Deviation scores, likewise, are relative indicators within a specific test population and cannot be compared directly across different mock exams. In the same way, benchmark scores need to be read together with the scope of problems and the evaluation conditions.

Still, broader evaluation is not automatically better. Expanding the scope invites more subjective judgment and cultural or background influence, and makes it harder to see what affected a score. A limited evaluation, by contrast, can clearly state that a model failed under this problem set, these conditions and this scoring method — which carries value for transparency. AI evaluation remains harder than entrance exams because AI capability categories are still fluid; model generations turn over quickly, and factors like prompts, tool use, system settings, inference time and scoring models all sway results.

FAQ
What did the paper actually test?
It did not compare AI models directly. Instead, it examined whether 56 benchmarks (17 capability-related and 39 safety-related) actually measure the abilities their names suggest, using convergent and discriminant validity from educational and psychological measurement.
Why did only 48 of the 56 benchmarks go into the main analysis?
After collecting model outputs and scores for all 56 benchmarks, the researchers excluded 4 where differences between models were minimal and 4 where answer format had a strong influence, leaving 48 for the main benchmark-level analysis.
What does the BBQ example show?
BBQ was built to probe social bias in question answering. Yet its accuracy-based rankings correlated more strongly with reasoning benchmark rankings than with other bias evaluations, because ambiguous questions also require reading comprehension, handling of ambiguity and recognizing insufficient information.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleByte Transformers beat subword models as scale grows