A stealth model hit Arena. Devs say it's Gemini 4 Pro
An unnamed checkpoint appeared on the Arena leaderboard this week, and the AI world has already decided what it is. The evidence is thinner than the excitement — which is exactly why the story is worth reading carefully.
A model carrying the name "gemini-3.8-flash" is competing anonymously on Arena, and developers who have been testing it say it is Google DeepMind's unreleased Gemini 4 Pro. Google has not confirmed the listing, has not commented on the speculation, and its own model documentation still shows Gemini 3.1 Pro as the newest Pro-tier model in the family — a generation it shipped in February, about seven months ago. The "stealth launch" framing is community inference, not an announcement, and it is worth keeping that line drawn: we covered what happened the last time Google's flagship slipped, when Brin took over Gemini as 3.5 Pro was reportedly shelved.
What the testers actually saw is more concrete than the identity question. An X post from developer Harshith showing a model-generated pelican-on-a-bicycle SVG drew 172,700 views and a wave of follow-up testing: 14-minute sketch-style landing pages, pixel-art 3D pagodas, an Airbus H145 model in about 10 minutes, working mini-games, and 3D flight simulations that clearly outperform the real, officially released Gemini 3.8 Flash. Responses reportedly take eight to ten minutes — long enough that the model appears to be spending heavy test-time compute planning and debugging before it answers. That last comparison is the only reliably falsifiable claim in circulation, and it is the one that keeps the rumor alive: whatever this is, it is not simply 3.8 Flash.
Then there is the benchmark chart. A single image circulating on X and reposted across Chinese tech outlets credits the mystery checkpoint with four category-leading scores — roughly 88 percent on DeepSWE v1.1, 2,064 Elo on GDPval-AA v2, 95.3 percent on Terminal-bench 2.1, and 86.8 percent on OSWorld-2.0 — against GPT-6 Astra and Claude Fable 5.1. It has no official source, no test configuration, and no reproduction link. Treat it as a screenshot, not a result; leaked eval charts have a habit of quietly failing independent reruns.
The backend claims deserve the same skepticism. Researchers who say they traced requests into the model report a 10-million-token input ceiling, a 256,000-token output ceiling, permanent cross-session memory, and web access without an API key — plus pricing of $2.25 per million input tokens and $11.25 per million output tokens. The posts describe that as the cheapest of the big three. It is not: Google's official Gemini 3.8 Flash lists at $0.75 and $3.75 per million tokens, so the leaked numbers are roughly three times the price of the model this checkpoint is allegedly impersonating. Flagship pricing, in other words, not a value play — and a good reminder that leaked pricing is usually wishful arithmetic.
The reason any of this moved markets of attention is scarcity. Gemini 3.5 Pro was previewed at I/O, missed its June window, and was reported cancelled this month by the Wall Street Journal on the grounds that it was not a clear enough step up from Flash. Into that vacuum, an unlabeled, obviously capable checkpoint reads as the flagship Google has been holding back, with an October launch the most commonly repeated — and entirely unconfirmed — expectation.
The recursive-self-improvement subplot is even further out on the limb. Some accounts claim Gemini 4 finished pretraining early because DeepMind closed an RSI loop internally; there is no paper, no statement, and no evidence attached. What is documented: DeepMind's chief strategy officer told a Berkeley audience that recursive self-improvement is central to the investment case for AI, and Google published research on September 14 on agents that improve their own search strategy across attempts. Both are real. Neither proves the loop closed. Z.ai made a narrower version of this argument public this week, which we covered in GLM's Infra Agent built the stack that serves GLM — and the gap between "an agent wrote part of our serving stack" and "a model improved itself" is where the whole debate lives.
What to watch: whether Google's API model list changes, whether any of the four benchmark scores shows up on an official leaderboard, and whether the Arena listing keeps the 3.8 Flash name after a real launch.
Rumors like this one get resolved by a single documentation page — so is patient skepticism the right posture, or does the rumor mill genuinely predict launches better than the labs do? Tell us in the comments.
Sources: 36kr — 刚刚,Gemini 4 Pro偷跑上线,碾压Astra和Fable · Qiniu Cloud — Gemini 4 Pro偷跑上线?匿名现身Arena · NokiaPowerUser — Gemini 4 Pro stealth-tested on Arena.ai · Google — Gemini API model list · X — Harshith on the Arena checkpoint