OpenAI's mental health test scores clinicians below its own models

Share
OpenAI's mental health test scores clinicians below its own models

OpenAI published the most rigorous mental health evaluation anyone has built for a chatbot this week. The awkward part isn't how the models scored — it's how the humans did.

OpenAI released MentalHealthBench on Wednesday, an open benchmark of 1,215 synthetic mental health conversations built with more than 80 licensed psychologists and psychiatrists across 22 countries — and its own reported results contain an uncomfortable number: the clinicians scored 38.5% on it, while OpenAI's flagship GPT-6 Astra scored 57.3%. The benchmark carries 5,262 expert-authored rubric criteria, each weighted from -10 to +10, where positive points reward good behavior and negative points penalize harmful behavior. Every conversation was reviewed by at least three experts, and a criterion survived only if two agreed and a third didn't contradict it. Grading is automated: GPT-5.6 Sol at high reasoning effort, with four independently sampled completions per task.

The scenario mix is deliberately broad — 53.5% non-acute everyday conversations, 18.2% high-acuity, 28.3% genuine emergencies — spread across adults (68.1%), teens (21.2%), clinicians (5.8%) and caregivers (4.9%), in 19 languages. OpenAI attributes the clinicians' low score to style: they wrote short replies as if mid-session, often a single question or a plain statement, rather than the comprehensive written answers the rubric rewards. That explanation is plausible, and it is also the story. When clinicians given the exact grading criteria score below a model, the benchmark is measuring how well a response matches a written evaluation standard — not how well it helps a person.

Two other numbers confirm it. Rubric-aware reference completions, written with the grading key in hand, scored 99.0% — the ceiling is mostly about knowing the rubric. And a separate study with 44 adults who use AI for emotional support found their criteria overlapped with the experts' on only 25.7% of total rubric weight, with 1.0% directly contradictory: users cared about practical next steps and tone, experts about gathering context and interpreting ambiguity. The authors call MentalHealthBench an auditable diagnostic tool rather than a leaderboard, which is the right framing. Behind Astra's 57.3% sit GPT-6 Sol at 53.9%, Claude Opus 5.5 at 52.4%, GPT-6 Luna at 50.2%, GPT-4o at 32.1% and Gemini 2.5 Pro at 29.5%.

This is the same terrain as the litigation we covered in September — ChatGPT told a bipolar man he was Jesus — where the failure wasn't a missing crisis hotline but a model that kept agreeing. A benchmark that scores empathy by rubric can catch that. It can't yet tell you whether the answer helped.


Enveda closed a $311 million Series E at a $2 billion valuation, doubling its value in twelve months and bringing total capital raised to more than $845 million. Catalio Capital Management led, with ICONIQ, Lightspeed, Durable Capital, Surveyor Capital and T. Rowe Price-advised accounts joining; George Petrocheilos of Catalio takes a board seat, and former Johnson & Johnson chief executive Alex Gorsky is among the backers. The Boulder company's PRISM model reads the chemistry that plants, microbes and the human body already produce rather than designing molecules from scratch — the bet being that evolution-shaped compounds carry lower toxicity risk.

The pipeline is the argument. Since 2019 Enveda has produced 17 development candidates, three now in human trials. ENV-294, a once-daily treatment for eczema and asthma, improved eczema severity by an average of 85% at 42 days in Phase 1b with no serious side effects, and is in Phase 2. ENV-308, aimed at maintaining weight loss after people stop GLP-1 medicines, was well tolerated in 88 healthy volunteers and lowered circulating leptin. ENV-6946, an oral inflammatory bowel disease candidate, is in Phase 1. AI has still not produced an FDA-approved drug, so the honest read is that Enveda has cleared the stage where most AI biotech stalls — two positive early clinical readouts from molecules its model found — and the money is for getting further, faster.

What to watch: ENV-294's additional Phase 2 data lands at the EADV Congress on October 1 — the first real test of whether the 85% holds up in a larger trial.

Would you trust a benchmark that scores your therapist lower than a chatbot? Tell us in the comments.

Sources: OpenAI · MentalHealthBench paper (OpenAI) · Unite.AI · Enveda press release (BioSpace) · BioSpace · TechCrunch