Nvidia's Vera Rubin posts a 30x agentic jump in first hardware tests

Share
Nvidia's Vera Rubin posts a 30x agentic jump in first hardware tests

Nvidia just put real numbers on its next flagship. At Hot Chips 2026 the company published the first silicon-level benchmarks for the Vera Rubin NVL72 rack, and it ran them on DeepSeek-V4-Pro — a 1.6-trillion-parameter open-weights model — doing agentic coding work rather than canned chatbot queries. Against the current GB300 NVL72, Nvidia claims up to 30 times more throughput per megawatt and per-token costs cut by as much as 35 times.

The comparison baseline matters as much as the headline number. Nvidia says GB300 already delivers roughly 15x H200 throughput per watt on the same workload, so Rubin's 30x lands on top of an already-steep curve — which is why some engineers reacted with disbelief ("I wouldn't even put that in a pitch deck"). These are vendor numbers from Nvidia's own developer blog, not independent measurements, so they should be read as a claim about where the Pareto frontier sits, not settled fact. Still, the direction is unambiguous: the benchmark that produced them, SemiAnalysis' AgentX, replays real coding sessions with growing context, tool calls, and sub-agents instead of fixed-length sequences. Nvidia's argument is that chat-era benchmarks are obsolete now that agents consume an order of magnitude more tokens than chat, so racks will be sold on tokens-per-megawatt of agent work — the metric that decides how much AI fits inside a data center's power budget.

The rest of the announcement rounds out the platform story. The new Vera CPU (88 custom cores, LPDDR5X at 1.2 TB/s) handles agent orchestration; SpaceXAI said it has moved to full-scale deployment of Vera and plans space-based "Starmind" satellites running Vera Rubin NVL72 racks from 2028. Groq 3 LPX — the low-latency inference accelerator Nvidia acquired in its $20 billion Groq deal — is in mass production, pairing with Rubin GPUs (GPU reads context, LPX writes tokens) and hitting a reported 3,400 tokens/second on Gemma 4 31B at 100K context. Nebius is already deploying it.

What to watch: whether any third party can reproduce the 30x figure on rented Rubin capacity — and what OpenAI's Jalapeño team quotes when its competing numbers land.

If a 35x drop in token cost holds, which part of your AI bill actually shrinks first — or does usage just expand to eat it? Tell us in the comments.

Sources: Nvidia Developer Blog — Vera Rubin and Blackwell set a new standard for agentic AI performance · Nvidia Developer Blog — Inside Nvidia Groq 3 LPX · Nvidia Developer Blog — DSX MaxLPS power management · AI Era (新智元) · Hot Chips 2026 coverage