AIToday
AI Safety & AlignmentHugging Face BlogPublished: Sep 23, 2026, 04:00 JST

UK AISI shares benchmark results via EvalEval's Evaluation Cards

UK AISI shares benchmark results via EvalEval's Evaluation Cards

3 Key Points

  1. What happened

    The UK AI Security Institute (AISI) is using EvalEval's Evaluation Cards to openly share verified evaluation methods and findings for five benchmarks—HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0—covering six frontier models including Claude Opus 4, GPT-5, and their updates.

  2. Why it matters

    This gives researchers and practitioners verified reference points to interpret evaluations and compare findings across the ecosystem, which is expected to support more reliable meta-research.

  3. What to watch

    The release accompanies AISI's paper on how inference compute and evaluation protocol affect benchmark performance, so its usefulness hinges on whether others also provide comparable setup details; watch for wider adoption of the Every Eval Ever schema.

WHO IT HITSAI evaluation researchers and policy analysts gain verified reference data to compare model performance across studies, reducing the risk of misinterpreting scores produced under different conditions.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

The collaboration between the UK AI Security Institute and the EvalEval Coalition began at a joint workshop alongside NeurIPS 2025, and feedback from the Institute helped shape the Every Eval Ever schema. This new phase moves that shared infrastructure into practice, with AISI releasing evaluation data through Evaluation Cards. The release includes results for five benchmarks and six frontier models, alongside a paper titled How Inference Compute Shapes Frontier LLM Evaluation, which examines how benchmark performance depends on inference-time compute and evaluation protocol. AISI has also worked on making evaluation more efficient through OptStop and more statistically rigorous through HiBayES.

The problem both groups are addressing is that evaluation results are reported across many formats and platforms, often without enough information to reproduce them, and re-running evaluations can be prohibitively expensive. By providing verified results, context, and configuration information, this release may offer a reference point for interpreting evaluations in context, for example by helping researchers understand how setup choices influence reported performance. As more evaluators adopt the Every Eval Ever schema, open comparisons like these could support broader meta-research.

The stakes hinge on whether other evaluators follow suit in sharing setup details, since the value of such reference points depends on broad participation. For now, researchers and policy analysts gain a concrete set of verified data to examine individual studies more closely and compare findings across the wider ecosystem.

FAQ
Which benchmarks are included in the release?
The release covers HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0. It also includes two related cyber evaluations, Cyber CTFs and The Last Ones.
What is Evaluation Cards?
Evaluation Cards is an open platform from the EvalEval Coalition that combines benchmark metadata, evaluation-run data, and model metadata into interpretable records.
Who can contribute to this effort?
Model developers, evaluation developers, and evaluation, governance, and policy researchers are invited to report verified results and use the Every Eval Ever schema or Evaluation Cards.
Hugging Face BlogRead Original Article

Get the latest AI Safety & Alignment news every morning

For example, today's edition would include:

  • Bessent: No government liability shield for AI labsFortune AI · 2h ago
  • Compute-to-GDP Fallacy skews AI pacing debateFortune AI · 2h ago
  • Ex-Indeed CEO blasts $140 million anti-regulation push before Trump-Xi AI talksFortune AI · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleOpenAI sets four priority areas for third-party safety audits