
What happened
The UK AI Security Institute (AISI) is using EvalEval's Evaluation Cards to openly share verified evaluation methods and findings for five benchmarks—HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0—covering six frontier models including Claude Opus 4, GPT-5, and their updates.
Why it matters
This gives researchers and practitioners verified reference points to interpret evaluations and compare findings across the ecosystem, which is expected to support more reliable meta-research.
What to watch
The release accompanies AISI's paper on how inference compute and evaluation protocol affect benchmark performance, so its usefulness hinges on whether others also provide comparable setup details; watch for wider adoption of the Every Eval Ever schema.
WHO IT HITSAI evaluation researchers and policy analysts gain verified reference data to compare model performance across studies, reducing the risk of misinterpreting scores produced under different conditions.
Summaries like this, in your inbox every morning.
The collaboration between the UK AI Security Institute and the EvalEval Coalition began at a joint workshop alongside NeurIPS 2025, and feedback from the Institute helped shape the Every Eval Ever schema. This new phase moves that shared infrastructure into practice, with AISI releasing evaluation data through Evaluation Cards. The release includes results for five benchmarks and six frontier models, alongside a paper titled How Inference Compute Shapes Frontier LLM Evaluation, which examines how benchmark performance depends on inference-time compute and evaluation protocol. AISI has also worked on making evaluation more efficient through OptStop and more statistically rigorous through HiBayES.
The problem both groups are addressing is that evaluation results are reported across many formats and platforms, often without enough information to reproduce them, and re-running evaluations can be prohibitively expensive. By providing verified results, context, and configuration information, this release may offer a reference point for interpreting evaluations in context, for example by helping researchers understand how setup choices influence reported performance. As more evaluators adopt the Every Eval Ever schema, open comparisons like these could support broader meta-research.
The stakes hinge on whether other evaluators follow suit in sharing setup details, since the value of such reference points depends on broad participation. For now, researchers and policy analysts gain a concrete set of verified data to examine individual studies more closely and compare findings across the wider ecosystem.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Treasury Secretary Scott Bessent told CNBC the government will not absolve AI labs of responsibility, saying t…

The authors argue both camps in the pacing debate have fallen for the "Compute-to-GDP Fallacy"—the belief that…

Former Indeed CEO Chris Hyams blasts accelerationism as the real AI threat ahead of a White House Trump-Xi sum…

Anthropic announced Tuesday that Claude Opus 5.5 is its "strongest-performing" model on its most comprehensive…

OpenAI proposed four priority areas for independent third-party assessment of frontier AI: safety cases, criti…

At Semafor's The Next 3 Billion event, Nvidia sustainability head Josh Parker attributed recent US anti-AI sen…
