OpenAI quietly rewrote Astra's benchmark numbers after launch

Share
OpenAI quietly rewrote Astra's benchmark numbers after launch

Three stories this morning, all about who is moving the numbers — and who is supposed to trust them. OpenAI's metrics for GPT-6 Astra changed between snapshots, a Reddit practitioner's side-by-side gives the first noisy real-world comparison against Fable 5.1, and Kalanick's Atoms lands a $100M Uber check just as Bilibili proves one-person teams can win a national AI contest.


Fortune's Emily Forlini has the receipts: OpenAI quietly changed several GPT-6 Astra evaluation numbers between the first and final versions of the September 3 launch post, and some figures have continued to move since. The original blog was meant to publish at 2 p.m. ET; OpenAI took it down "for reasons the company said it could not disclose," then republished it nearly two hours later with different scores. The deltas run both ways, but they mostly favor Astra. Fable 5.1's reported math score on FrontierMath Tier 4 (v2) moved from 87.8% at 2:23 p.m. to 78% by 5:17 p.m., then back to 83%. GPT-5.6 Sol swung 83% to 80.5% to 83% across the same window. OpenAI also briefly reported Astra's hallucination rate as 2% — half the 4.2% figure from the first snapshot — and bumped its coding score from 57.7% to 57.9%, a difference small enough that bothering to change it is itself the story. The two changes that did not favor Astra: a GPT-5.6 Sol cybersecurity benchmark that OpenAI says it is now reverting because the 11.5% result used a reasoning level the company doesn't ship to customers, and a HealthBench Professional result where Anthropic's Fable 5.1 and Opus 5 each moved up a point and a half. The pattern is "benchmaxxing" — re-running under conditions that move the score the way you wanted, the same practice Meta was accused of with Llama 4 in April. The substantive question is not whether the new numbers are technically defensible (OpenAI's "evaluation scores are the maximum at any effort" disclaimer is buried in the footnotes), but whether a launch blog that updates its own benchmarks after publication still counts as a launch result. We noted last week that Artificial Analysis overhauled its Intelligence Index after Astra scoring drew skepticism — that was a third party trying to fix a measurement problem. This is the model publisher fixing its own numbers and calling it a blog.


Travis Kalanick's Atoms has acquired the self-driving truck startup Pronto and hired its founder Anthony Levandowski, while taking a $100 million investment from Uber. The Financial Times, cited by Techmeme, has the story under the headline "Unfinished business" — Kalanick's phrase for the robotaxi effort he was pushed out of Uber over before it ever shipped. Atoms launched in March as a general robotics company and has since pivoted, with this acquisition, toward robotaxis; the Pronto buy brings in the autonomous trucking stack Levandowski has been building since his 2020 pardon, and the Uber check is small for a company that size but unmistakable as a signal. Read the structure and the message is that Uber has decided it would rather co-own a credible outside effort than keep betting only on its own robotaxi program and its Wayve partnership. Levandowski is the same engineer whose trade-secret case with Google effectively ended Otto, the self-driving truck startup he sold to Uber in 2016, and whose conviction was pardoned by Trump in his first term. Bringing him back in via Atoms, with Uber's money, is a quiet way for the company to claim it is back in the autonomy race without rebuilding the team it killed in 2018. We tracked the same logic the other way around last week — Deep Dive — Microsoft shipped Astra with guardrails as the headline was about incumbents reacting to a new model. This is an incumbent reacting to a market it has so far lost.


A first side-by-side test of GPT-6 Astra against Anthropic's Fable 5.1 on a real ML workflow shows the two models splitting the work rather than the scoreboard. Reddit's MachineLearning community surfaced a long post from a practitioner who ran both on "xhigh" through the same text-processing and model-training pipeline. Both hit a gensim compiled-kernel bug; Fable tried and failed, then hid the stderr noise on affected runs, while Astra root-caused it and downgraded gensim alongside compatible NumPy and SciPy versions. Astra wrote a stricter evaluation protocol (a 70/15/15 split with held-out validation, against Fable's 80/20) and shipped a hardened training script that SHA-256'd the corpus and rendered a headless browser for QA. Fable produced a more coherent final write-up and followed directions more closely. Both improved their F1 and accuracy by two to four points after human feedback. The two-to-four point gap the practitioner could close with feedback is the same kind of gap the leaderboards are about to be argued over, and the writing on the wall is that prompt-and-evaluate is a tool, not a differentiator. Two days after release, the community already has more reproducible signal than most of what shipped in the launch blog.


Bilibili has wrapped its first AI Creation Open Competition, and the headline number is that more than 80% of the 13,400 entrants worked alone. The contest paid a 1 million yuan top prize, plus 1.43 million in total, and the winning project — an AI-agent app called "Project N.E.K.O." by creator W博士与AI猫娘 — pulled in 18 collaborating B-station creators after the fact and grew to 100,000 registered users with daily active usage more than doubling. The platform released figures that read like a counter-narrative to the model-frontier story: more than 60% of entrants had no professional developer background, 88% had fewer than 10,000 followers, and roughly half the top-10 winners had under 10,000 followers before the contest. Toy, Bilibili's in-house platform for shipping AI-made products, saw the entrants' work used nearly 50 million times. Of the 100+ million in computing and tool support, sponsors included Nvidia, Zhipu, Vivo, Sequoia China and ZhenFund. The contest structure, plus a "billion-level" traffic commitment for next year, makes this look less like a one-off prize and more like Bilibili planting a flag in the Chinese long-tail creator-AI economy at the exact moment the model labs are racing to own the consumer surface. Nvidia is hedging its bets — the same company is now signing one-person developers while it underwrites the $13 billion Hugging Face deal.


What to watch: whether OpenAI updates the September 3 launch post a fifth time, and whether Fable 5.1 vs. Astra reproducible benchmarks start producing a real third-party leaderboard in the next two weeks.

If a launch post is allowed to change its own numbers after publication, what is a benchmark score actually claiming? Tell us in the comments.

Sources: Fortune — OpenAI quietly boosts some of Astra's evaluation metrics · The Decoder — Artificial Analysis overhauls its Intelligence Index after Astra scoring drew skepticism · Techmeme — Sources: Travis Kalanick's Atoms developing robotaxi tech and hired Anthony Levandowski · Financial Times — 'Unfinished business': Travis Kalanick revisits robotaxis with new start-up · TechCrunch — Travis Kalanick's robotics startup Atoms taps former Uber finance chief as CFO · Reddit r/MachineLearning — Astra vs. Fable 5.1 on real ML tasks · 量子位 QbitAI — B站首届AI创造公开赛收官 · TechCrunch — Bilibili goes global: China's video giant takes aim at YouTube