China Mobile Cloud runs attention on GPUs and FFN on brain chips
Two announcements out of China this week point the same direction: when the chip node is capped, the wins have to come from system architecture — and from a build-out policy that finally measures utilization instead of volume.
China's first production system that splits a large model across domestic GPUs and brain-inspired chips went live today, and the vendor claims it more than doubles inference output while cutting operating cost by more than 40%. Mobile Cloud — China Mobile's cloud arm — built the heterogeneous inference system with the Nanhu Research Institute of China Electronics Technology Group, Beijing-based neuromorphic chipmaker LingXi Technology, Shanghai GPU vendor Iluvatar CoreX, Tsinghua and Peking University. The mechanism is the interesting part: the team performs a PD/AF separation of the Transformer, sending prefill and attention computation to the domestic GPU and the feed-forward network to the neuromorphic chip, whose compute-in-memory design and large on-die SRAM are a natural fit for the bandwidth-hungry half of the layer. On top of that sits a self-developed model compiler, a high-speed interconnect protocol and a unified inference engine that splits tasks, schedules across both chip classes and aggregates results. Validation ran on DeepSeek V4 Flash: against a same-budget all-GPU domestic cluster, inference output and energy efficiency both improved by more than 1× and business operating cost fell more than 40%, per the project team. Fifteen granted invention patents and six software copyrights are claimed, with small-batch trial production already under way, and Mobile Cloud says the core heterogeneous framework will be opened up gradually to domestic GPU and neuromorphic vendors and compute operators. The stated benchmark is Nvidia's next-generation Vera Rubin-plus-Groq heterogeneous inference stack — the disaggregated-prefill playbook, rebuilt on chips China can actually buy. That is the honest read: this is an architecture workaround for a process-node ceiling, not a silicon win. The numbers are vendor-supplied and single-model, and the earlier evidence on domestic clusters was humbler — we covered the scale problem in DeepSeek plans a 160,000-chip Huawei cluster in Inner Mongolia and the price problem in China's AI chip prices jump 50% as the memory shortage bites. Splitting a layer across two chip architectures is how you get a fourth of the bill down without getting better transistors.
Beijing's five-year roadmap for the telecom sector stops treating compute scale as the goal and starts treating compute structure as the problem. In the 15th Five-Year Plan for information and communications technology, laid out at a State Council Information Office briefing this week, MIIT official Xie Cun reported China's intelligent computing capacity at 2,185 EFLOPS as of the end of June 2026, distributed 55.9% east, 32.6% west, 10.6% central, 0.9% northeast, with an overall rack turn-on rate of 71.4%. The plan calls for a hub–region–edge hierarchy, "orderly" deployment of 10,000-card and 100,000-card-plus clusters, and inference capacity deployed on demand by scenario rather than by headline FLOPS — plus automated compute monitoring to match supply with actual use. The most concrete market consequence is thermal: the plan explicitly steers new facilities toward liquid cooling. TrendForce puts liquid cooling at roughly 33% penetration on AI chips in 2025 and forecasts 53% for 2026, because single-chip TDP now routinely exceeds 1 kW and full AI racks run to hundreds of kilowatts; even so, fewer than 40% of Chinese AI data centers are liquid-cooled today. Cold-plate designs lead in the near term, with immersion following as rack density climbs. What to watch: whether "orderly" becomes a permit. A plan that rewards utilization over installation is a direct insult to the GPU-count arms race — the question is whether any province can be stopped from buying cards.
If splitting a model across two kinds of mediocre chips beats buying one kind of excellent one, is that innovation or an indictment of export controls? Tell us in the comments.
Sources: CNRI (央广网) · Kuai Technology via NetEase · China Times via QQ News · MIIT, interpretation of the 15th Five-Year Plan for ICT · Guangming Online