Google's Gemini 4 Argon gets a 1M-token output limit

Share
Google's Gemini 4 Argon gets a 1M-token output limit

Gemini 4 Argon's output ceiling jumps to one million tokens. Google confirmed in its launch post that Argon can now generate up to one million output tokens, up from the 64K limit on prior Gemini models — an industry first at that scale. Pricing lands at an introductory $2 per million input tokens and $10 per million output, with cached input 95% off, rising later to $4 and $20. The catch: Argon isn't broadly available yet — Google is rolling it out first to a small group of trusted cyber defenders under its Fairwind program, with wider API access undated. Independent testing tempers the headline numbers: Artificial Analysis puts Argon at 53 on its Intelligence Index, tying GPT-6 Astra but trailing Claude Opus 5.5's 58, and notes it burns roughly 62,000 output tokens per task versus Astra's 27,000. A one-million-token ceiling only pays off if the model knows when to stop writing — right now the interesting spec is the price, not the limit. We covered the internal reaction yesterday — Google employees question Gemini 4 Argon's real-world coding.


OpenAI's enterprise billing now covers Chinese open models — including Moonshot's Kimi K3. Baseten announced a partnership that puts it among the first open-model inference providers in OpenAI's B2B Marketplace, with a native integration inside Codex: enterprise customers can spend their existing OpenAI commitments on served open models like Kimi K3 and GLM 5.3, via Codex or the Responses API, on a waitlist for now. Chinese outlets are running it as a first — the first Chinese model inside OpenAI's enterprise payment system — and the timing is the real story. OpenAI's own blog post yesterday attributed a large-scale model-distillation campaign to individuals associated with Moonshot AI, so the same week OpenAI publicly blamed Moonshot-linked accounts for stealing reasoning traces, it opened a path for Moonshot's flagship to be bought with OpenAI purchase orders. We covered that attribution on Tuesday — OpenAI pins its reasoning-theft campaign on Moonshot AI accounts.


Flow raises 50 million dollars to point AI agents at hardware engineering. Flow Engineering, which sells software agents that watch CAD files, Git repos, simulations and requirement docs to catch conflicts and failed requirements before they reach the shop floor, closed a $50 million Series B at a $750 million valuation, co-led by Valor's Antonio Gracias and Atreides' Gavin Baker, with Sequoia's Roelof Botha joining the board. The traction claim that matters: Rivian went from 40 to 1,500 internal users in seven months, and the customer roster reads like a defense-and-aerospace index — Anduril, Joby, Stoke Space, GM's powertrain unit. It's the "agents for the physical world" thesis without the robots: the hardware is the customers', Flow just designs alongside it.

What to watch: whether Google puts a date on Argon's general API access, and how many enterprises actually spend OpenAI dollars on Kimi K3.

Is a one-million-token output limit a feature or an invitation to ramble? Tell us in the comments.

Sources: Google blog · The Decoder · Baseten · Shanghai Metals Market · TechCrunch · Unite.AI