Google's Gemini 4 Carbon reportedly matches Opus 5.5 on coding

Share
Google's Gemini 4 Carbon reportedly matches Opus 5.5 on coding

Google hasn't launched Gemini 4 Argon yet, but internal documents suggest a faster follow-up is already in testing — and early impressions put it level with Anthropic's best coding model.

Google is testing a Gemini 4 variant called "Carbon" that employees say performs on par with Anthropic's Opus 5.5 on programming tasks. Business Insider, citing internal documents, screenshots, and chats, reports that Google deployed Carbon on Jetski — its internal coding platform — over the past few days, with at least one employee comparing its coding ability to Opus 5.5, Anthropic's strongest model. Carbon is one of several internal Gemini 4 variants alongside Argon (the announced frontier model) and Barium, whose Barium-B checkpoint was selected as the public Argon release. It likely ships as an update within the Argon family rather than a separate tier; one employee called it internally the "Gemini pro next model."

The timing is what makes this more than a leak. Argon only just arrived, and Carbon's arrival in internal testing within days suggests Google's model iteration loop is accelerating. DeepMind employee Vedant Misra responded to the report on X with "Have you heard of recursive self improvement" — a nod to the idea that AI is increasingly building better AI. OpenAI and Anthropic have both reported similar dynamics in their own pipelines, and Google recently mapped its lineup explicitly: Argon for frontier reasoning, Flash for speed, Omni for media, Gemma for edge. A Carbon drop would slot straight into that frontier slot before Argon even reaches general availability. We covered Argon's coding reputation before — Google employees question Gemini 4 Argon's real-world coding — and Carbon reads like the direct answer to those internal doubts.

Worth keeping the report proportionate: the Opus 5.5 comparison comes from one employee and, by Business Insider's own account, still needs more testing. Google has not commented. What to watch: whether Carbon ships as an Argon update or stands alone, and whether any of it shows up on public benchmarks before the Gemini 4 launch date lands.

Is AI building the next generation of AI models faster than labs can name them — or is one employee's chat message being over-read? Tell us in the comments.

Read more

Only 4.5% of US consumers pay for AI, and the top 1% spends $900

Only 4.5% of US consumers pay for AI, and the top 1% spends $900

Consumer AI has near-half household reach and a paying base the size of a rounding error — and today's numbers on who actually spends put a fine point on it. Only 4.5% of US consumers pay for AI, and the top 1% of them spends $900 a month. Andreessen Horowitz published the seventh edition of its Top 100 generative AI apps ranking, and for the first time the firm tracked observed spending on US consumer cards alongside traffic and downloads. The picture: nearly half of US adults now use AI, abo

Open Source Radar — October 10: reviewers, skills, and 3D maps

Open Source Radar — October 10: reviewers, skills, and 3D maps

Today's trending board skews practical: tooling that fixes code, teaches agents a framework, and rebuilds 3D scenes from video. Here's what's worth your stars. alibaba/open-code-review (Go, 45.6k stars) — Alibaba open-sourced the code review tool it runs internally, and it shows: a hybrid of deterministic static-analysis pipelines for the known bug classes (null dereferences, thread-safety, XSS, SQL injection) plus an LLM agent that writes precise, line-level comments. It talks to OpenAI- or A

Deep Dive — Congress has the data center numbers, still no bill

Deep Dive — Congress has the data center numbers, still no bill

A yearlong Senate investigation into seven of the biggest data center developers in the country concluded this week that the public case for the AI buildout does not survive the companies' own paperwork — and it landed at the exact moment Congress needs it, because the one federal bill written to make data centers pay for their own power fell three votes short of advancing in the Senate, despite passing the House 417-3. The report, led by the offices of Senators Elizabeth Warren, Chris Van Holle

WorldArena 2.0 puts world models to the test on real robots

WorldArena 2.0 puts world models to the test on real robots

The world-model field has argued for two years about video quality. This week the first full results landed for a benchmark that asks the harder question: can a robot actually use the prediction? WorldArena 2.0's global challenge has closed its leaderboard, and for the first time world models are being graded on physical robots instead of plausible video. The benchmark — designed by a Tsinghua-led consortium with PKU, CMU, Stanford, Princeton and others — extends its 1.0 video scoring along th