Story
OpenAI Releases Benchmark to Evaluate AI Model Responses in Mental Health Scenarios

Summary
OpenAI has launched MentalHealthBench, an open-source framework developed with licensed experts to assess the quality and safety of AI-generated responses in mental health conversations. The company also published initial scores for several large language models against the new standard.
OpenAI on Wednesday released MentalHealthBench, a new open-source benchmark designed to evaluate the performance of AI models in realistic mental health conversations. The framework was developed in collaboration with more than 80 licensed mental health experts from 22 countries to standardize the assessment of AI capabilities in this sensitive domain.
Benchmark Structure and Performance
MentalHealthBench assesses AI models on key behaviors such as safety, seeking context, preserving user agency, and providing actionable guidance. According to OpenAI, the benchmark is composed of various scenarios:
- 53.5% involve non-acute, everyday conversations.
- 28.3% simulate emergencies with immediate safety concerns.
- 18.2% cover high-acuity conversations that indicate serious mental health issues.
OpenAI evaluated several AI models against the benchmark, with its own GPT-5.6 Sol model serving as an automated grader against expert-defined criteria. The company reported that GPT-6 Astra achieved the highest score at 57.3%, followed by GPT-6 Sol at 53.9% and Claude Opus 5.5 at 52.4%.
Expert-Led Evaluation Criteria
AdThe evaluation system was built upon detailed rubrics created by the participating mental health experts, with each scenario reviewed by at least three professionals. Criteria were assigned weights from -10 to +10 based on clinical importance to score the model responses.
In a separate analysis involving 44 adults who had used AI for mental health support, OpenAI found that users valued practical next steps and an appropriate tone. In contrast, clinical experts placed a higher emphasis on the AI's ability to gather relevant context and correctly interpret ambiguous situations. The final scoring criteria for the benchmark were based on the expert consensus.
Context and Broader Safety Measures
The release of MentalHealthBench is part of a broader effort by OpenAI to enhance safety in its products. The company stated it has also strengthened ChatGPT’s responses in sensitive conversations, expanded access to crisis resources, and introduced new features like a "Trusted Contact" option and a specialized "ChatGPT for Teens" with additional protections. By making the benchmark open to researchers, OpenAI aims to foster further development and scrutiny in the field of AI for mental health.
Read next
More on Stocks
McDonald's to Launch Tiered Loyalty Program to Boost Customer Frequency
McDonald's announced plans to introduce a tiered loyalty system to better differentiate rewards for its most frequent customers and drive repeat sales. The move, detailed at the company's investor day, aims to create more personalized experiences and follows a broader industry trend away from margin-eroding discounts.

Trump Administration to Fast-Track Vape and Nicotine Pouch Approvals, WSJ Reports
The Food and Drug Administration is preparing to ease regulatory requirements for smoke-free nicotine products, a move that could significantly benefit major tobacco companies but faces opposition from public health advocates, according to The Wall Street Journal.

Peloton Reportedly Explores $800 Million Debut Bond Sale
Peloton Interactive is considering its first-ever bond offering to raise $800 million, according to a Bloomberg report citing people familiar with the matter. The news sent the company's shares down over 4% in Wednesday trading.

1789 Capital Seeks $3 Billion for Second Growth Fund, Bloomberg Reports
The investment firm, where Donald Trump Jr. is a partner, has reportedly secured $2 billion toward its target for a new fund focused on larger investments in sectors like AI and defense.