
Anthropic Fellows built TASTE, a benchmark for AI safety research judging. Models agree with human experts only 60% of the time.
Humans agree with each other about 77% of the time.
This suggests AI judges are not yet reliable.
What happened
Anthropic Fellows built TASTE, a benchmark of 92 AI safety research proposal pairs. Human experts picked better proposals; models agreed with them only 60% of the time, worse than humans.
Why it matters
AI safety progress often lacks clear right answers. A reliable AI judge could help evaluate research cheaply, but current models fall short of expert judgment.
What to watch
The benchmark's design uses expert discussion and strong-confidence labels to reach 77% estimated agreement. Future models may improve on TASTE, but no such result is reported yet.
Ask the AI about this article →
TASTE was created as part of the Anthropic Fellows Program, addressing the challenge that many AI safety research questions cannot be evaluated with verifiable rewards. The benchmark uses pairs of proposals and aims for high human agreement through discussion and confidence filtering. The 60% model performance indicates that current AI models are not yet reliable judges for this kind of research. This underscores the difficulty of automating evaluation in fields where expert judgment is still essential. Future work might focus on improving model judgment, but no such results are reported in this article.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Dozens of protesters gathered Thursday outside 8VC, a venture capital firm in Austin, to protest Palantir's AI…

ECRI, a nonprofit patient safety organization, is expanding its reporting system to include errors involving a…

Visa says its AI 'harness' makes Anthropic cheaper to use for cyber defense

Visa announced enhancements to its cybersecurity portfolio, including the next evolution of its open-source Vi…

OpenAI, Anthropic, Microsoft, Alphabet, Amazon, and over 100 other companies issued a joint letter Thursday ca…

More than a hundred registered nurses protested outside Palantir's former Palo Alto headquarters on Aug
