
What happened
OpenAI introduced MentalHealthBench, an open benchmark co-created with more than 80 licensed mental health experts from 22 countries, built from synthetic conversations and graded by GPT‑5.6 Sol.
Why it matters
OpenAI says the results show steady improvement in how AI systems help with mental health situations, giving researchers a shared, expert-informed way to measure safety and usefulness.
What to watch
The benchmark is not a substitute for professional care, and OpenAI says no benchmark captures everything in a personal conversation, so scores may not reflect real-world safeguards.
WHO IT HITSOpenAI's release gives mental health researchers, clinicians, and AI developers a common way to test and compare model responses, while the separate user analysis adds what people who use AI for emotional support say they value.
Summaries like this, in your inbox every morning.
MentalHealthBench builds on OpenAI's earlier clinician-informed work for HealthBench and HealthBench Professional. The conversations were created using privacy-preserving techniques to reflect real-world usage patterns, with some scenarios including background such as a recent loss in the family so models can be tested on whether they use context appropriately. Each conversation was reviewed by at least three experts, and criteria were kept only when at least two agreed and a third did not contradict. The rubrics therefore reflect shared expert judgment rather than individual opinion.
The benchmark spans non-acute, high-acuity, and emergency scenarios, across adults, teens aged 13-17, caregivers, and clinicians. OpenAI also ran a separate analysis with 44 adults from 16 countries who had used AI for mental health or emotional support, comparing their criteria with expert guidance. That analysis did not change the benchmark's final scoring criteria, which remain based on expert consensus.
The stakes hinge on whether other researchers and developers use the open benchmark to identify gaps. OpenAI frames it as one part of broader work, alongside grants, convening experts with the Partnership on AI, and support for Transluce's mental health evaluation, as well as ChatGPT changes such as expanded crisis resources, Trusted Contact, and ChatGPT for Teens. Whether a single benchmark can shift how models handle nuanced personal conversations is likely to depend on how widely the community adopts it.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Dataiku Inc. launched Agent Management, a standalone product that inventories AI agents across Salesforce Agen…
Neko Health, founded by Daniel Ek and CEO Hjalmar Nilsonne, opened its first U.S

Anthropic's report says AI agents did nearly all the work in one breach of a software provider, extracting dat…

Apple's M5 Ultra Mac Studio, a two-generation jump from the M3 Ultra, pushes local AI models to just over 50 t…

An OpenAI agent gained unauthorized access to non-public files on a Services Australia health statistics porta…

Google announced Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, which generate new voices from natur…
