Nvidia's Groq 3 LPX inference chip enters full production

Share
Nvidia's Groq 3 LPX inference chip enters full production

Nvidia commercialized its largest acquisition at Hot Chips 2026 on Monday, while a New York startup betting that video games can teach robots how to move pulled in fresh capital at a $6 billion valuation. Here's what moved in AI hardware and physical AI.


Nvidia said its dedicated inference accelerator Groq 3 LPX has entered full production, turning the company's record $20 billion purchase of Groq into a shipping product. The chip is built as an extension of the Vera Rubin data center platform, offloading the token-generation phase of inference so agents can reason, call tools and iterate without the lag that makes them feel sluggish. Nvidia packages 256 of the accelerators into a single rack, and in an Artificial Analysis benchmark the rack hit a record 3,400 tokens per second running an open Gemma 4 agentic model across a 100,000-token context window — roughly four times more responsive on latency-sensitive workloads than rival platforms, per the company. Nebius signed on as the first customer, planning to put the racks to work in its Nebius Token Factory; Groq racks will be online later this year. Nvidia also announced that SpaceXAI will build its next-generation AI architecture around Vera Rubin, adopting the Vera CPUs for orchestration across data centers and orbital satellites. Nvidia reports earnings Wednesday. Low-latency inference is turning into the industry's polite arms race — OpenAI's new "Ultrafast" mode, powered by Cerebras, promises 750 tokens per second, and AMD is pairing its systems with Cerebras racks too. The twist here is that Nvidia is using that $20 billion Groq deal to own the decode phase, betting the crown jewel of the AI era isn't training scale but how fast a model appears to think.


Valor Equity Partners, Point72 and Seven Seven Six are backing General Intuition at a $6 billion pre-money valuation as the startup raises fresh funds to push into robotics. The New York company, spun out of gameplay-clip platform Medal last October, builds a foundation model that trains AI agents to move through space and time, using hundreds of millions of hours of video games plus their "action labels" as its initial dataset. Existing investors Khosla Ventures and General Catalyst are joining the round, which comes weeks after General Intuition raised $320 million at a $2.3 billion valuation; a person close to the deal says it is oversubscribed, and the funds will go toward compute (General Intuition partners with CoreWeave) and talent as it focuses on robotic embodiments. Valor's rare backing — its first AI lab since SpaceX — is a signal of how hot the physical-AI bet has become. The premise is that embodied intelligence can be seeded not in the physical world, where data is scarce and slow, but in the digital replay of human action, then generalized to robots. It is one of the more audacious wagers in AI, and the market is pricing it accordingly.

What to watch: whether Groq 3 LPX's token-generation speed shows up in benchmarks that people actually feel in chatty coding agents — and whether Nvidia's Wednesday earnings confirm the raucous demand the company keeps signaling.

If inference speed is the new battleground, does the winner need to build its own chip — or is renting from the cloud enough? Tell us in the comments.

Sources: SiliconANGLE · CNBC · Techmeme · TechCrunch · NVIDIA