AIToday
AI Safety & AlignmentLessWrong AIPublished: Aug 29, 2026, 06:00 JST1 min read

New TASTE benchmark tests AI judges on safety research

New TASTE benchmark tests AI judges on safety research

Key takeaway

  • Anthropic Fellows built TASTE, a benchmark for AI safety research judging. Models agree with human experts only 60% of the time.

  • Humans agree with each other about 77% of the time.

  • This suggests AI judges are not yet reliable.

3 Key Points

  1. What happened

    Anthropic Fellows built TASTE, a benchmark of 92 AI safety research proposal pairs. Human experts picked better proposals; models agreed with them only 60% of the time, worse than humans.

  2. Why it matters

    AI safety progress often lacks clear right answers. A reliable AI judge could help evaluate research cheaply, but current models fall short of expert judgment.

  3. What to watch

    The benchmark's design uses expert discussion and strong-confidence labels to reach 77% estimated agreement. Future models may improve on TASTE, but no such result is reported yet.

Ask the AI about this article →

Context & Analysis

TASTE was created as part of the Anthropic Fellows Program, addressing the challenge that many AI safety research questions cannot be evaluated with verifiable rewards. The benchmark uses pairs of proposals and aims for high human agreement through discussion and confidence filtering. The 60% model performance indicates that current AI models are not yet reliable judges for this kind of research. This underscores the difficulty of automating evaluation in fields where expert judgment is still essential. Future work might focus on improving model judgment, but no such results are reported in this article.

FAQ

What does TASTE measure?
TASTE measures how well AI models can judge pairs of AI safety research proposals, scored by agreement with experienced human researchers' preferences.
How well did models perform on TASTE?
Models achieved 60% agreement with human researchers, which is lower than the human agreement level of 77%.
What design choices helped build TASTE?
Including a discussion stage where researchers talk through disagreements and filtering labels for self-reported strong confidence helped achieve a high-agreement benchmark.

Get the latest AI Safety & Alignment news every morning

For example, today's edition would include:

  • Nurses protest Palantir's AI in hospitalsTop Companies AI · 4h ago
  • ECRI expands reporting to cover AI errorsTop Companies AI · 4h ago
  • Visa expands AI cybersecurity tools to fix threats fasterTop Companies AI · 4h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleLambda secures $1B debt for Nvidia chips