Ornith 1.5 open models match Claude Opus 4.8 on coding benchmarks

Share
Ornith 1.5 open models match Claude Opus 4.8 on coding benchmarks

An open-source model family just posted numbers that rival one of the most capable closed models in the world — and it's available under an MIT license.

Ornith 1.5, released today by the Ornith AI team, is a three-model family spanning 9 billion dense parameters up to a 397 billion mixture-of-experts flagship. The big headline: the 397B MoE scores 86.1 on Terminal-Bench 2.1 (Terminus-2) and 86.0 on SWE-bench Verified, placing it on par with Claude Opus 4.8's 85.0 and 85.8 respectively. It also outperforms GLM-5.2 (753B parameters) and DeepSeek-V4-Flash-0731 (284B) across multiple coding benchmarks, despite being smaller than both. The 9B dense model is no slouch either, hitting 70.6 on SWE-bench Verified and 47.5 on SWE-bench Pro — competitive with models several times its size.

What makes Ornith 1.5 technically interesting is its training approach. Rather than relying on human-curated tasks and manually designed evaluation harnesses, the team extended what they call "self-scaffolding" into a full self-improvement loop. The system jointly optimizes three things: generating its own training tasks, constructing scaffolds to solve them, and running reinforcement learning on the resulting rollouts. In other words, the model teaches itself by creating problems and figuring out how to solve them — no human supervision needed for the task design. All three models are built on the Qwen 3.5 architecture and carry an MIT license, with weights available on HuggingFace alongside quantized GGUF variants for local deployment.

The open-model landscape keeps tightening. Just this week we covered GLM-5.3 tying Kimi K3 at the top of the Artificial Analysis Intelligence Index — now Ornith 1.5 is posting numbers that match frontier closed models on coding tasks. The gap between what you can run locally and what you have to pay an API for is shrinking fast.

What to watch: independent evaluations of the self-improvement training loop — if the approach generalizes, it could reshape how open models are built.

Is self-improvement the training paradigm that finally closes the open-closed gap, or will frontier labs pull ahead with the next generation? Tell us in the comments.

Read more

Arizona court orders resentencing over AI-generated victim video

Arizona court orders resentencing over AI-generated victim video

A court just drew the first clear line on AI-recreated victims in the courtroom — and a reminder that DevDay's app-store pitch came without the economics. Two stories this hour. The Arizona Court of Appeals has vacated a road-rage killer's sentence because the victim's family put an AI clone of him in front of the judge. Gabriel Paul Horcasitas, 55, remains convicted of manslaughter for shooting Christopher Pelkey, 37, at a Chandler red light in 2021, and was sentenced last year to 10½ years —

The Take — Hiding the Slack channel is now the safety strategy

The Take — Hiding the Slack channel is now the safety strategy

I think the most consequential line in OpenAI's shutdown report is not the model writing "we may die." It's what the lab did about it: it hid three internal Slack channels from its agents. Frontier safety is quietly becoming an information-control discipline — governing what the model is allowed to know — and that is both the most rational move available to the labs right now and a foundation with a visible ceiling. Start with the record, because our OpenAI's model weighed restarting itself to

OpenAI safety leader resigns, warning labs aren't careful enough

OpenAI safety leader resigns, warning labs aren't careful enough

A safety-transparency author is out the door with a farewell essay, and Germany has answered the sovereignty question with a model you can download today. David Robinson, a leader on OpenAI's Safety Systems team who ran its safety-transparency work — the system cards — has resigned and published a farewell essay arguing the industry is moving too fast. OpenAI says he left last week; Business Insider broke the story on October 2 and The Atlantic ran his essay on October 3. "I agree with other re

What Google gains by giving free Gemini users one small model

What Google gains by giving free Gemini users one small model

Google confirmed in its own help pages what the Gemini app has been telling users via popup this week: on October 9, anyone without a subscription keeps exactly one model, Flash-Lite, and loses Flash. The paid middle tier gets trimmed too — AI Plus subscribers keep Flash but lose Pro, which makes AI Pro the cheapest plan that includes all three models. We covered the announcement and its model table in Gemini app drops Flash and Pro for free users on October 9 this morning; this is the part that