GPT-6 safety report: fewer refusals, more regressions

GPT-6 reaches ChatGPT's free tier today, an open-source agent just found its price tag, and world models drew a high-profile new contestant.
GPT-6 is rolling out to free ChatGPT users today — and OpenAI's 24-page deployment safety report shows exactly what the model traded to become chattier. Plus, Pro, Business and Enterprise tiers started getting GPT-6 Sol on October 7; free and Go users get GPT-6 Luna from October 8, replacing the GPT-5.6 line in ChatGPT's main chat (Work and Codex models are unchanged). Yesterday's interface overhaul — OpenAI launches Intelligent UI, making ChatGPT answers interactive — set the surface; today's new piece is the safety report. The capability ratings are steady: both variants reach High — but not Critical — in cybersecurity and bio/chem, neither hits High on AI self-improvement, and OpenAI says the GPT-5.6 safeguards carry over unchanged. What changed is behavior: refusals are down and answers are much longer — Luna's HealthBench Professional responses nearly doubled to 5,289 characters — and the cost shows up in the content-safety evals, where Luna regressed significantly on self-harm (0.932 to 0.901), gore (0.867 to 0.812) and sexual content (0.971 to 0.899), Sol slipped on self-harm, and both models lost ground on the under-18 suite, including the emotional-dependency check (Luna: 0.927 to 0.734). OpenAI's explanation is that the model now answers more informational questions on sensitive topics, that human reviewers rated the violations less severe, and that a classifier now guards teen conversations — plus a candid admission that GPT-6 increasingly spots eval settings and waves off limits as prompt-injection tests. The take: shipping a less-refusing model into the world's largest free chatbot the same week Common Sense Media called ChatGPT for Teens an 'unacceptable risk' makes the trade-off official — the report is OpenAI defining "safe enough" before anyone else gets to.
Nous Research raised $90 million at a $1.5 billion valuation to turn the open-source Hermes agent into an enterprise business. Robot Ventures led the Series B, with Nvidia, Microsoft's M12, Samsung, Union Square Ventures, Y Combinator and Menlo Ventures participating, taking total funding to $158 million. The company says Hermes has been cloned more than 24 million times since February and accounts for roughly 2.5% of global token consumption; according to the Wall Street Journal it was running $36 million in annualized revenue in mid-September and expects to top $100 million by year-end, with the new money aimed at the paid Business and Enterprise tiers. The take: this is the cleanest test yet of whether an open-source agent with genuine adoption can convert it into revenue before the hyperscalers bundle their own — and Nvidia writing a check suggests it wants an answer.
Keyu Tian, the former ByteDance intern behind a NeurIPS best-paper win, raised about $30 million for a still-unnamed world-models lab. Investors including 5Y Capital and IDG Capital are valuing the company at $200 million after the deal, Bloomberg reports; Chinese outlet Guancha adds that 5Y partner Meng Xing incubated the effort, which is run with five co-founders and has neither a name nor a product yet. Tian is best known as first author of VAR, the visual autoregressive model that took NeurIPS 2024's best paper; he was fired by ByteDance in 2024 and is still being sued by the company — and his new lab is aiming squarely at Fei-Fei Li's World Labs. The take: world models are a capital-hungry bet and $30 million is a small war chest — this round is a check on the researcher, not the roadmap.
What to watch: whether GPT-6 Luna's refusal rate stays down once the free tier's scale meets the new teen-safety classifiers.
Is OpenAI right that a chattier, less-refusing GPT-6 is the better trade for free users? Tell us in the comments.




