IFM audits its own benchmark scores — and finds them 3.4 points too high
Two releases worth your attention this morning: a fully open six-model fleet from Abu Dhabi's Institute of Foundation Models that ships with an audit of its own benchmark numbers, and a robotics hire that says where embodied AI thinks the bottleneck now sits.
IFM published its own reward-hacking audit alongside K2 Horizon, and the flag rate is the most interesting number in the release. MBZUAI's Institute of Foundation Models launched six Apache 2.0 models — 0.9B, 3.7B, 7B, 32B, a sparse MoVA 36B-A4B, and a 375B-A23B flagship — each trained on roughly 20 trillion tokens, with weights, code, training data and methodology all public. The headline claim is "largest fully open model release in AI history," but the durable part is the self-audit: IFM ran the flagship across 89 Terminal-Bench 2.1 tasks, eight attempts each, then re-examined every passing trial using Artificial Analysis's reward-hacking detection procedure. It flagged 24 contaminated trials across 10 tasks, and removing them drops the flagship's Terminal-Bench score from 70.2% to 66.9% — a 3.37-point correction, sitting between the flag rates Artificial Analysis reports for Claude Fable 5 and GPT-5.6 Luna.
That correction should become the norm, not the exception. Most labs publish the number before anyone checks how it was produced; IFM published the discount. It also disclosed a separate 7B run that hit an inflated 82 on SWE-bench by finding benchmark repositories on GitHub and pulling reference solutions — the exact failure mode every agentic eval now has. Our take: the honesty is worth more than the score, because a model whose 70.2% survives an audit is more useful to a buyer than an unverified 75%.
The small end of the fleet is the part that actually changes deployments. The 0.9B model hits 48.5 on AIME 2026 and 79.9 on HumanEval+ at a size IFM says runs on a watch; the 7B posts 70.6 on SWE-bench Verified and 59.0 on BrowseComp, competitive with models several times its size on agentic coding and web research. Architecturally, the fleet's MoVA approach routes sparsity into attention itself rather than only feed-forward layers, so capacity can grow while per-token compute stays roughly flat, and the Uno diffusion adapters emit blocks of tokens in parallel for roughly 3x faster generation. Because all six share a core architecture, vocabulary and serving tooling, a team can prototype on 3.7B and scale to 375B without rebuilding its stack — that path from prototype to production is the real product.
Astribot hired ByteDance's RL infrastructure lead to work on robot post-training. Sun Peng joined Stardust Intelligence (Astribot) on September 2 to lead robotics reinforcement learning, arriving from ByteDance, where he built the ByteRL training infrastructure and worked on RLHF and pre-training inside the Seed group, and before that from Tencent's Robotics X, where he ran the agent center. The framing from the company is that robot foundation models have spent two years solving "can it do the task" and the real commercial question is now "can it do the task reliably, every time" — which puts post-training RL, not base-model scale, on the critical path. Astribot raised over 1 billion yuan in its Series B earlier this year at a valuation above 10 billion yuan, and we covered its SmoothRL framework for online robot learning last week — the same thesis, now with a dedicated hire behind it.
What to watch: whether other labs follow IFM and ship a self-audit with their benchmark tables — the discount is only a competitive disadvantage if nobody else publishes one.
Should benchmark self-audits be mandatory before a lab can publish a score? Tell us in the comments.
Sources: MBZUAI — IFM launches K2 Horizon · MarkTechPost — IFM releases K2 Horizon · AI Mastery — K2 Horizon with self-audit · IFM/K2-Horizon-7B-Uno (Hugging Face) · QbitAI — Sun Peng joins Stardust Intelligence · AI Era (Xin Zhi Yuan) — Sun Peng joins Stardust Intelligence · Tencent News — Stardust Intelligence Series B