Chinese AI labs stay on Nvidia as CUDA lock-in holds

Share
Chinese AI labs stay on Nvidia as CUDA lock-in holds

The buildout's fault lines are showing: China's top labs are still training on Nvidia despite the export squeeze, Goldman puts global AI investment past $1 trillion this year — and KPMG finds nearly half of executives are pulling back AI agents because the bills beat the benefits.

China's most advanced AI models are still being trained on Nvidia chips, and software switching costs — not the hardware ban — are the reason, sources at major Chinese LLM developers tell the South China Morning Post. Beijing's push for self-sufficiency was supposed to move frontier training onto Huawei's Ascend line, but its CANN programming environment demands that developers rewrite and re-optimize large amounts of CUDA code, trading away Nvidia's finely tuned libraries for performance uncertainty at exactly the moment labs are racing to ship competitive models. It's a reminder that the domestic-hardware milestone we covered earlier today — China's first fully domestic 100,000-card supercluster — solved the supply problem, not the software lock-in. Huawei has fought back by open-sourcing CANN and embedding migration engineers at customer sites like Baidu and Tencent, but CUDA's moat is network effects: every developer, library, and benchmark built around it makes the next switch harder, not easier.


Goldman Sachs Research now expects global AI investment to top $1 trillion in 2026 — roughly $581 billion of it in the US. The bank's economists built the number by augmenting the commonly cited $794 billion US hyperscaler capex forecast, which they argue understates global AI spending by around $200 billion while overstating the US share by a similar amount. Add private companies, non-US players, and AI-exposed firms beyond the hyperscalers, and cumulative AI investment reaches $1.8 trillion by year-end, with AI capex climbing from 1.8% of US GDP this year toward 2.8% by 2028 — levels consistent with prior general-purpose technology buildouts.


Nearly half of executives — 49% — have scaled back AI agent deployments because operating costs outweigh the benefits, according to KPMG's Q2 2026 Global AI Pulse survey of 2,145 senior leaders across 20 countries. The pullback isn't a retreat: 79% still call AI a top investment priority, up from 74%, with average planned spend holding at $188 million. The real problem is metering — agents burn tokens on every tool call and self-check, and most companies can't see the meter: only 26% report real-time visibility into what AI costs to run, and a third cite limited understanding of token pricing as a barrier to deployment. It's the same cost squeeze we flagged when SAP froze most travel and hiring as AI costs soared — SAP freezes most travel and hiring as AI costs soar — and the emerging fix is financial, not technical: cost dashboards, token literacy, and cost reviews baked into approval loops before agents scale.

What to watch: whether Huawei's CANN open-sourcing and on-site migration teams start bending the switching-cost curve — and whether enterprise agent budgets stabilize or keep contracting.

If your company runs AI agents, do you actually know what the token bill looks like? Tell us in the comments.

Read more

OpenAI busts influence ops that planted fake stories in real media

OpenAI busts influence ops that planted fake stories in real media

The day's AI news runs through one seam: the work is showing up in places nobody planned for — inside real newsrooms, across the whole night sky, and in the M&A column. OpenAI has banned two state-backed influence operations that used ChatGPT to plant fabricated stories inside legitimate news outlets — and rated the Russian one the most disruptive it has seen in two and a half years. In a report dated October 8, OpenAI detailed "Dark Clark," run from Russia across Latin America, which ran a th

Open Source Radar — October 9: plugins, sandboxes, tokens

Open Source Radar — October 9: plugins, sandboxes, tokens

Today's open-source signal is infrastructure rather than hype: Microsoft's code sandbox reaches 1.0, Anthropic's knowledge-worker plugins keep climbing, a beloved token counter flips its default, and LocalLLaMA squeezes a usable 2B model into about 700 MB. knowledge-work-plugins (Python, ~27,900 stars, Apache-2.0) — Anthropic's repository of role-shaped plugins for Claude Cowork is the top AI repository on today's daily trending page, and the stars keep coming: roughly 2,100 more than when we

Deep Dive — The four-token blind spot inside DeepSeek V4

Deep Dive — The four-token blind spot inside DeepSeek V4

ByteDance's Seed research team says it has found the cause of one of the stranger recurring complaints about DeepSeek's models: the same question, asked with nothing changed except a few junk characters bolted onto the front, can flip the model from right to wrong. Their paper, posted to arXiv on September 28, traces the wobble to a memory-saving trick used during long-context inference, and reports that DeepSeek-V4-Flash-Base's retrieval accuracy swings by as much as 40.2 percentage points depe

SoftBank seeks $100B from Gulf investors for an AI fund

SoftBank seeks $100B from Gulf investors for an AI fund

Three moves today point the same direction: the money, the politics, and the price of speed all got more expensive. SoftBank is reportedly seeking up to $100 billion from Gulf investors for a fund that would buy companies and run them with AI. The Financial Times reported the raise, citing people familiar with the matter, and says Masayoshi Son has held discussions in recent weeks with senior figures including in the United Arab Emirates; Reuters and Bloomberg both carried the report but neith