Qualcomm's next NPU runs 30B-parameter models on a phone
Two very different bets on where AI actually runs landed today: Qualcomm says the next Snapdragon will carry a 30-billion-parameter model in your pocket, and China's industry ministry picked the car as the first industry where AI models and agents get to run at scale.
Qualcomm has detailed its next-generation Hexagon NPU, and the headline number is 30 billion parameters on a phone. The company says mixture-of-experts models up to 30B total parameters will run on the NPU, with roughly 3 billion active parameters routed per token — the MoE trick of carrying a big model's knowledge while only paying for a slice of it each step. The NPU gets a new Element Accelerator for transformer math, vector and scalar extensions that Qualcomm pitches for agent decision-making and routing, support for context windows up to 32K, and hardware KV-cache acceleration. Shared memory grows by 50 percent over the Snapdragon 8 Elite Gen 5, though Qualcomm still won't say what the absolute capacity is. For INT4 models it claims up to 50 percent faster prefill — the part where the model chews through your prompt before it starts answering.
The interesting part is not the parameter count, it's what Qualcomm thinks the NPU is now for. The scalar and vector extensions exist to handle the bookkeeping of an agent — deciding what to call, routing between steps — not just multiplying matrices. That is a company designing silicon for the assumption that the thing running in your hand will be a loop of tool calls rather than a single prompt and answer. The catch, as always: these are Qualcomm's own numbers against its own prior generation, and no shipping phone has been benchmarked. Treat 30B as a ceiling for a demo, not a promise about battery life.
China's industry ministry and eight other departments told the auto industry it goes first. The nine departments' "15th Five-Year" plan for intelligent connected new-energy vehicles instructs the sector to push innovation in AI models and agents and apply them in cars ahead of other industries, including speeding up construction of pilot-scale testing bases for automotive AI applications and encouraging firms to build compute capacity through a mix of public and self-owned facilities. The practical read: compute for training, inference and scenario validation may shift from every company building its own toward a blended model, and cars — dense with sensors, cameras and driving data — become the reference deployment for agentic AI in Chinese industry.
This is the same ministry that has been setting hard targets for AI service providers, and the pattern is consistent: name an industry, build the compute, mandate the pilots. Putting agents in vehicles raises the stakes on a different axis than a chatbot does — an agent that misroutes a request is annoying, one that misroutes a decision in traffic is a recall. The plan's emphasis on testing bases suggests Beijing knows that, and wants a place to fail before the fleet does. We covered the ministry's broader AI application playbook in China's industry ministry lays out its AI application playbook for the next five years.
OpenAI's $1-a-year government deal is over, replaced by half-price metered tokens. The General Services Administration announced a new 27-month OneGov agreement giving agencies a 50 percent discount on token-based usage across ChatGPT models starting Oct. 1 — the day after the current deal expires. There is no platform fee, no minimum order and no spend commitment; the standard licence is $15 per user per month. OpenAI says the agreement extends eligibility to roughly 23 million people, against more than one million government employees who have access today, and every verified government entity gets its advanced cyber-defender system Daybreak Blue at half of standard commercial pricing. GSA puts OneGov's total savings at about $1.68 billion, of which $1.4 billion comes from AI agreements.
Free was never going to last, but the timing is sharper than it looks: Anthropic's and Google's initial OneGov deals also expire at the end of the month, and GSA has signalled it wants to renew. So the three largest US model providers are repricing their relationship with the federal government in the same week — from pilots priced to be ignored to contracts priced to be used. Agencies that built workflows on a dollar will now have to decide whether the work is worth a metered bill, which is arguably the point: consumption pricing is how a vendor finds out what you actually value.
What to watch: whether Qualcomm's MoE claims survive a shipping handset, and whether Anthropic and Google match OpenAI's 50 percent or hold out for better terms.
If your agency's AI budget went from $1 to a metered bill overnight, what gets switched off first? Tell us in the comments.
Sources: Qualcomm — Hexagon NPU: A new mobile architecture for agentic AI · ServeTheHome · HotHardware · Nextgov/FCW — GSA unveils new, token-based OneGov discount with OpenAI · OpenAI — Expanding AI access to the US government · 钛媒体 (Google News 中文) · 36氪 via 读懂AI时代