AIToday
Large Language ModelsAI in HealthcareAI Safety & AlignmentOpenAI BlogPublished: Sep 24, 2026, 06:00 JST

OpenAI launches MentalHealthBench, 80+ experts, 22 countries

OpenAI launches MentalHealthBench, 80+ experts, 22 countries

3 Key Points

  1. What happened

    OpenAI introduced MentalHealthBench, an open benchmark co-created with more than 80 licensed mental health experts from 22 countries, built from synthetic conversations and graded by GPT‑5.6 Sol.

  2. Why it matters

    OpenAI says the results show steady improvement in how AI systems help with mental health situations, giving researchers a shared, expert-informed way to measure safety and usefulness.

  3. What to watch

    The benchmark is not a substitute for professional care, and OpenAI says no benchmark captures everything in a personal conversation, so scores may not reflect real-world safeguards.

WHO IT HITSOpenAI's release gives mental health researchers, clinicians, and AI developers a common way to test and compare model responses, while the separate user analysis adds what people who use AI for emotional support say they value.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

MentalHealthBench builds on OpenAI's earlier clinician-informed work for HealthBench and HealthBench Professional. The conversations were created using privacy-preserving techniques to reflect real-world usage patterns, with some scenarios including background such as a recent loss in the family so models can be tested on whether they use context appropriately. Each conversation was reviewed by at least three experts, and criteria were kept only when at least two agreed and a third did not contradict. The rubrics therefore reflect shared expert judgment rather than individual opinion.

The benchmark spans non-acute, high-acuity, and emergency scenarios, across adults, teens aged 13-17, caregivers, and clinicians. OpenAI also ran a separate analysis with 44 adults from 16 countries who had used AI for mental health or emotional support, comparing their criteria with expert guidance. That analysis did not change the benchmark's final scoring criteria, which remain based on expert consensus.

The stakes hinge on whether other researchers and developers use the open benchmark to identify gaps. OpenAI frames it as one part of broader work, alongside grants, convening experts with the Partnership on AI, and support for Transluce's mental health evaluation, as well as ChatGPT changes such as expanded crisis resources, Trusted Contact, and ChatGPT for Teens. Whether a single benchmark can shift how models handle nuanced personal conversations is likely to depend on how widely the community adopts it.

FAQ
Who created MentalHealthBench?
It was co-created with more than 80 licensed mental health experts from 22 countries, including psychologists and psychiatrists speaking 19 languages.
How are model responses graded?
An automated grader, GPT‑5.6 Sol, evaluates responses against expert-written criteria, each weighted from -10 to +10 based on clinical importance.
Did users and experts agree on what is helpful?
Users highlighted practical next steps and tone, while experts placed greater emphasis on gathering context and interpreting ambiguous situations.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Dataiku launches Agent Management for cross-vendor AI agentsSiliconANGLE AI · 14h ago
  • Anthropic: AI agents hit roughly 200 firms' data in one breachFortune AI · 14h ago
  • M5 Ultra Mac Studio hits 50 tokens/sec as prices jumpArs Technica AI · 14h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleRabbit unveils OS3, a cloud AI agent for up to five machines