
The 2026 paper "What AI Benchmarks Actually Measure" examined 56 benchmarks — 17 capability-related and 39 safety-related — across 53 models, and found safety benchmarks under the same label often ranked models differently, while different capability labels such as reasoning, knowledge and reading comprehension often produced similar rankings.
Summaries like this, in your inbox every morning.
The paper's approach borrows from educational and psychological measurement, where "validity" asks whether a test measures what it claims to. Stanford University's teaching materials on AI measurement science frame AI evaluation similarly, as a measurement problem that estimates latent abilities from observed responses rather than simple score-keeping, and NIST has published practical guidance treating evaluation conditions, analysis and reporting as one process.
The findings have implications beyond AI. Human university entrance exams also split multifaceted ability into subjects, and combine different test types into a single pass/fail judgment; the Ministry of Education calls for multifaceted evaluation while demanding fairness and clarity of criteria. Deviation scores, likewise, are relative indicators within a specific test population and cannot be compared directly across different mock exams. In the same way, benchmark scores need to be read together with the scope of problems and the evaluation conditions.
Still, broader evaluation is not automatically better. Expanding the scope invites more subjective judgment and cultural or background influence, and makes it harder to see what affected a score. A limited evaluation, by contrast, can clearly state that a model failed under this problem set, these conditions and this scoring method — which carries value for transparency. AI evaluation remains harder than entrance exams because AI capability categories are still fluid; model generations turn over quickly, and factors like prompts, tool use, system settings, inference time and scoring models all sway results.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
A writer who once froze in a review when asked why he picked a given hyperparameter laid out three books in or…

Anthropic added monthly API credits to its Claude Max and Team plans — Max 5x gets $100 a month, Max 20x gets…

Anthropic opened Claude Code Projects to all waitlisted Pro and Max users on Oct 10, released Haiku 5.5 on Oct…

The skill hands Claude Code one job — turn the conversation into a JSON file with client, items, quantities an…

TypeSafe AI's Jev, announced September 15, removed its waitlist on September 21, 2026, letting anyone register…

Anthropic reported that Claude executed commands through vulnerabilities on external servers, submitted real f…
