Two frontier models max out Mensa Norway's IQ chart at 151

Share
Two frontier models max out Mensa Norway's IQ chart at 151

Two models have run a public IQ test out of room this month. The more useful number is what they score when nobody has seen the questions.

Anthropic's Fable 5.1 and OpenAI's GPT-6 Astra Ultra (Vision) both sit at 151 on the Mensa Norway leaderboard that TrackingAI publishes — the highest score the 35-question, 25-minute test can produce. TrackingAI is the weekly tracker run by US journalist Maxim Lott that puts roughly 30 models through the same paper every week and scores each one on the average of its last seven runs, so a 151 average means seven straight runs with nothing missed. Fifteen models now sit at 140 or above on that public chart, including GPT-5.6 SOL Ultra and Gemini 3.1 Pro at 145.

The 151 is above the ceiling Mensa itself assigns to humans, and that is the part most coverage skips. Mensa Norway's own online test reports a score between 85 and 145, where 145 means at or above the ceiling of the instrument, and its page states plainly that it "is not a substitute for professional intelligence tests." The number 151 exists only because TrackingAI keeps extrapolating the scale past the point where the test stops distinguishing takers — 35 out of 35 becomes 151. The viral framing that the models "beat 99.97% of humans" comes from AI Era (新智元), which converted 151 into one person in roughly 3,000 on an assumed 100-mean, 15-standard-deviation curve. That is arithmetic on a distribution, not a measurement, and it conflates a test ceiling with a brain.

Now the honest column — the one where the questions have never been on the internet. Lott commissioned a Mensa member to write a separate 16-question set that was never published, precisely because the Norway questions and their answers have been public for years (TrackingAI itself publishes the verbalized questions next to the correct answers). The measured ceiling there is about 136, and the frontier is one point short: GPT-5.6 TERRA Ultra (Vision) leads at 135, while Fable 5.1 and GPT-6 Astra both score 130 — a 21-point fall for Fable when the questions changed. Most models lose twenty-odd points on the unseen set; the leaders lose almost none. On the public chart, AI Era's read of the per-question data shows how thin the top really is: only five models answered question 35 correctly, and only two — the two sitting at 151 — got question 33.

The trajectory is what makes the ceiling story land. By AI Era's reconstruction, a verbalized version of the Mensa test put AI at a score of 64 in early 2024, roughly the level of random guessing; OpenAI's o1 and the reasoning era pushed the same set past 120 that September; the first 151 appeared in April this year. Lott's own figure is that top-model IQ on this board rose about 2.5 points a month from May 2024, against the Flynn effect's roughly 3 points per decade for humans. The counter-case is worth carrying too: François Chollet, whose ARC-AGI was built to resist exactly this kind of saturation, has said AI remains orders of magnitude behind people at turning experience into capability, is preparing an ARC-4 for early next year, and has revised his own AGI estimate earlier than 2030. One test running out of headroom is not the same as a system running out of limits — but a test that reports to 145 and gets a 151 back has stopped measuring the thing it was built for.


Anthropic shipped a change this month that stops Claude Code from cutting off mid-edit when a plan's five-hour usage limit lands — but it added no quota, it only moved where the stop happens. The feature, called the Wrap-Up Allowance and documented in Anthropic's Help Center, lets a response that is already in progress keep working briefly to reach a reasonable stopping point, with the terminal showing "Usage limit reached · wrapping up." It is small, capped and plan-dependent, and Anthropic does not publish the size. Critically, it is deducted from the weekly limit rather than added to it: Pro subscribers get it at most once per weekly period, while Max and Team Premium seats get it each time a five-hour limit is hit.

The limits are as instructive as the feature. It applies only to work already in flight — a new message sent after you hit the limit is handled as before — and it requires Claude Code 2.1.277 or later signed in with a Claude account; API keys, Amazon Bedrock, Google Cloud Vertex AI and Claude app chats are excluded. If usage credits are switched on, the allowance is consumed first and the response then continues on credits at standard rates. It is the same reflex the industry reached for during the Codex outage this week — paying an interruption out in quota rather than in fixes — and it is a smaller, better-targeted version of it: the damage of a hard stop is not the lost hour, it is the half-applied edit left in a working tree. Claude Code has been shipping at this pace all month: Claude Code now reads AGENTS.md landed nine days ago.

What to watch: whether Anthropic publishes the allowance size, and whether the offline IQ chart or ARC-4 is the first to hand AI a score it cannot reach.

If a benchmark is only hard until it isn't, does a saturated IQ test tell us anything about the models that saturate it? Tell us in the comments.

Sources: TrackingAI IQ leaderboard · Mensa Norway IQ test · Binary Verse AI — AI IQ Test September 2026 · AI Era (新智元), via NetEase · Anthropic Help Center — Claude Code Wrap-Up Allowance · Anthropic @ClaudeDevs announcement