Anthropic runs an internal 'Model 2' stronger than any public Claude

Share
Anthropic runs an internal 'Model 2' stronger than any public Claude

Two stories today point at the same quiet truth: the most interesting AI isn't always the version you can log into. Anthropic is running a model more capable than any Claude you've met — and a freshly funded startup is betting that video's next form is code, not pixels.


Anthropic is quietly running an unreleased model internally that beats every public version of Claude. According to the company's August 2026 Risk Report, the model — codenamed "Model 2" and placed in the Mythos class — scores about 1.5 points above Claude Mythos 5 on Anthropic's internal capability index (AECI), though the gain is smaller than the jump from Opus 4.6 to Mythos. Anthropic leans on it heavily for coding, data generation, and research, including agents that run continuously, and says Claude now writes most of the code in its production systems. The company rates misalignment risk as "low," notes the model wasn't tested as thoroughly as Mythos 5, and has no plans to release it externally. The telling detail isn't the benchmark — it's that the most capable model a leading lab operates isn't the one users can touch. When the internal model already authors most of the lab's own code, the public frontier becomes a lagging indicator of what these companies can actually do.


SigmaZ AI has raised several million dollars to build real-time interactive video from diffusion language models. The startup, backed by BlueRun Ventures, Yunqi Capital (an early Y Transformers investor), Dezon Investment and others, trains models that emit executable frontend code rather than a fixed pixel stream, so the resulting "video" can be edited live — change a price, a chart, or a 3D object and the frame updates without regenerating the whole clip. Its flagship product, Tap8, is slated to launch this year; a July demo drew tens of thousands of waitlist signups and more than 20 enterprise conversations within 24 hours. SigmaZ claims its engine topped Fable 5, Seedance and Veo on quality and factual coverage in an internal 173-clip benchmark, with an automated evaluator independently reproducing the ranking at Spearman ρ = 0.89. Code-as-medium is a genuinely different bet than the pixel-video labs — and a faster one, since diffusion language models can decode in parallel at over 2,000 tokens per second versus roughly 100 for autoregressive code generation. Whether consumers actually want "living" video is the open question, but the engineering thesis is sound.

Should the real "frontier" be measured by the model behind the API, or the one a lab keeps for itself? Tell us in the comments.

Sources: The Decoder · Anthropic Risk Report (PDF) · AI Era (Xin Zhi Yuan)