AI forecasters underestimated progress on every benchmark that resolved
The people best placed to call AI's next milestone keep calling it too late. An audit of five years of expert forecasts puts a number on the gap, and shows the error pointing the same way almost every time.
On four AI benchmarks that have already resolved, the domain experts surveyed by the Forecasting Research Institute had assigned an average probability of 24.6 percent to the outcomes that actually happened. The professional superforecasters did worse — 9.7 percent.
The interim report, published on September 23, aggregates forecasts the institute has collected since mid-2022, including its 339-person LEAP panel: 76 computer scientists, 76 AI industry figures, 68 economists, 119 policy researchers, and twelve people on Time's AI 100 list. The widest miss is mathematics. AI reached gold-medal level at the International Mathematical Olympiad in July 2025, five years earlier than the median expert forecast and ten years earlier than the median superforecaster's. On the Navier–Stokes problem — the result OpenAI published this month, and which the Clay Mathematics Institute has not certified — experts had put 10 percent odds on AI settling a Millennium Prize problem by the end of 2027; superforecasters put 5.4 percent.
The pattern holds outside competitions. In AI virology, experts predicted models would not match a top virologist team on a troubleshooting benchmark until 2030 (superforecasters: 2034); the institute now dates that to roughly April 2025. On LiveCodeBench Pro's hardest tier, the median expert forecast 14 percent state-of-the-art performance by the end of this year and superforecasters said 12 percent — it was 53.8 percent by May 2026. On the METR task-horizon benchmark, the two groups forecast 3.4 and 3.5 hours by December; Anthropic's Mythos Preview was already at 3.1 hours in May.
Money is where the gap gets widest. Asked for the highest annual recurring revenue of any AI company at the end of 2026, experts said $20 billion, economists $16 billion, and superforecasters $25 billion — against roughly $100 billion reported for Anthropic and $40 billion for OpenAI on September 22. A newer question on the pair's combined run rate got medians of $70 billion and $90 billion against a current figure near $140 billion. The institute is candid about the limits: this is an interim analysis, overestimates only become visible once a deadline passes, so it is tilted toward finding caution, and some resolutions lean on model-generated projections of data the original forecasters never had.
The counterexamples matter for anyone tempted to generalise. Biosecurity experts predicted 22.5 percent of participants would complete lab tasks with a language model; only 5.2 percent did, no better than internet access alone. And forecasters may have been too optimistic about driverless cars. What has clearly changed is the forecasters themselves: their average odds on AI being a "technology of the century" rose from 31 to 36 percent for experts, and 28 to 35 percent for superforecasters, in nine months.
The take: the useful reading is not "experts are useless" — it is that the reference class everyone plans against sits years behind the observed curve. The same panel underestimated revenue by a factor of five. Institutional timelines built on expert consensus — procurement, safety review, regulation — are being set from forecasts that reality keeps beating, which cuts both ways: it argues against the plateau thesis and in favour of building slack into the slower-moving parts of the system.
What to watch: whether the next round of LEAP forecasts, taken after these misses, actually closes the gap or just shifts up by a quarter.
PrismML's 1-bit Bonsai language model now runs on Qualcomm's Snapdragon AR1 Gen 1 platform, which Qualcomm demonstrated at its Snapdragon Summit as the kind of model that can live locally on smart glasses. The version here is a 2-billion-parameter model tuned for vision and language, so the wearer can ask what they are looking at and get an answer without a round trip to a server. PrismML claims its compression keeps almost all of the larger model's benchmark performance while shrinking it fourfold — the same pitch behind the 27B Bonsai that fits on a laptop. No glasses shipping with it have been announced.
Apple researchers published a recipe for training speech recognition on users' devices without ever collecting their audio, improving on the strongest prior on-device method by 20.8 percent on average in-domain and 10.0 percent across domains. The paper's practical finding is that the two design choices can't be separated: letting each device learn from a copy of its own evolving model beats a single server-side teacher, but only if the server also keeps training on its small set of labelled data between rounds — otherwise the device model drifts. The paper landed on arXiv on September 21 with Apple's own research page listing it as a September 2026 publication — a research result, not a shipping feature.
Forecasts keep undershooting reality — so what should the rest of us actually plan on? Tell us in the comments.
Sources: Forecasting Research Institute — How Accurate Have AI Progress Forecasts Been So Far? · The Decoder · Axios — Anthropic's $100 billion revenue · Bloomberg — OpenAI's revenue run rate tops $40 billion · TechCrunch — PrismML brings its tiny LLMs to Qualcomm-powered smart glasses · PrismML · Apple Machine Learning Research — A Practical Recipe for Semi-Supervised Federated ASR · arXiv — Semi-Supervised Federated ASR