OpenAI Launches MentalHealthBench, an Open Benchmark for AI‑Driven Mental‑Health Conversations
OpenAI unveiled MentalHealthBench on September 23 2026, a publicly available benchmark created with over 80 licensed mental‑health professionals to evaluate AI responses across everyday, high‑acuity, and emergency mental‑health scenarios.
OpenAI Introduces a New Standard for AI Mental‑Health Support
On September 23 2026, OpenAI announced the release of MentalHealthBench, an open‑source benchmark designed to measure how AI systems respond in realistic mental‑health conversations. The benchmark was co‑created with a global cohort of more than 80 licensed psychologists and psychiatrists spanning 22 countries, speaking 19 languages, and representing nearly 20 mental‑health subspecialties. OpenAI says the tool will let researchers and developers assess model behavior on safety, context‑seeking, user‑agency preservation, and actionable guidance.
How MentalHealthBench Works and What It Reveals
MentalHealthBench evaluates AI across three acuity levels—non‑acute everyday talks, high‑acuity situations indicating significant distress, and emergencies that require immediate real‑world support. Scenarios cover adults, teens (13‑17), caregivers, and clinicians, and are generated using privacy‑preserving synthetic conversations that mirror real‑world usage patterns, sometimes including background details such as a recent family loss.
Each synthetic dialogue is reviewed by at least three experts; criteria that receive agreement from at least two experts and no contradiction from a third become part of the final rubric. The rubrics assign weighted scores from ‑10 to +10, rewarding beneficial behaviors (e.g., asking clarifying questions, offering empathetic reflections) and penalizing harmful ones (e.g., speculation, unsafe advice). An automated grader, GPT‑5.6 Sol, scores model responses against these expert‑written criteria.
OpenAI evaluated a range of contemporary models, including the newest offering from each provider as of the release date. Results show a steady improvement in handling mental‑health conversations, though OpenAI stresses that ChatGPT remains a supplement—not a replacement—for professional therapy.
In a parallel user study, 44 adults from 16 countries and 14 languages rated model replies to non‑acute synthetic conversations. Participants highlighted the importance of practical next steps and tone, while experts placed more weight on gathering context and interpreting ambiguous cues. The study did not alter the benchmark’s expert‑derived scoring but offered insight into user preferences.
Quotes underscore the initiative’s intent:
“Mental health exists on a continuum, from flourishing to everyday stress to acute crisis. AI systems that engage people across that range need to be grounded in both clinical science and lived experience—not only to recognize where someone falls on the continuum, but to know how to respond appropriately at any point.” – Dr. Arthur Evans, CEO of the American Psychological Association
“As a practicing psychiatrist, I often hear how patients use tools like ChatGPT to better understand their mental health and engage more deeply in their care. Bringing my clinical experience to this work lets me contribute beyond the hospital—helping ensure this technology empowers people while treating mental health with the care and responsibility it deserves.” – Dr. Kevin La Moureaux
OpenAI is also extending its mental‑health safeguards: expanded crisis‑resource links, a Trusted Contact feature for real‑world support, and a dedicated ChatGPT for Teens product with additional protections. The company notes that MentalHealthBench builds on earlier clinician‑informed efforts such as HealthBench and HealthBench Professional, and it will be offered openly for the community to examine, improve, and extend.
Industry Implications and Collaborative Efforts
By publishing MentalHealthBench, OpenAI joins a growing movement toward transparent, shared evaluation tools for AI safety and well‑being. The benchmark complements other initiatives, including grants for new mental‑health research, collaborations with the Partnership on AI, and support for independent projects like Transluce’s mental‑health evaluation framework. OpenAI’s open release invites competitors and academic groups to benchmark their models, potentially raising the overall standard for AI‑mediated mental‑health support across the industry.

