Deep Dive — Alibaba wants to own the whole stack. Memory says no.
The most important number Alibaba put on stage in Hangzhou this week was not ten trillion.
Alibaba opened its annual Apsara Conference on Tuesday with a plan that is unusually easy to restate, because the company said it in one sentence: it will train Qwen models of five trillion to ten trillion parameters, run them on its own accelerators wired into clusters of up to 500,000 cards, and operate more than 20 gigawatts of data centre capacity worldwide by 2032. Chief executive Eddie Wu gave the keynote framing — machines currently produce "less than 3 per cent" of all human thinking, and if that volume scales to a thousand times human capacity there is still an enormous runway — and chairman Joe Tsai supplied the theme, "Intelligence Goes Beyond."
Strip the framing and a strategy is legible in the three announcements taken together: own the model, own the silicon it runs on, own the power and the buildings. That is a specific answer to a specific problem. Nvidia's top-end parts are export-controlled out of China, TSMC cannot be used at the leading edge, and every one of those dependencies is a lever a foreign government can pull. If Alibaba owns the checkpoint, the accelerator, the network fabric and the power purchase agreement, then no licence, no entity list and no tariff sits between the company and the tokens it sells. As we covered this morning in Alibaba plans a 10T-parameter model and a 500,000-chip cluster, Broadcom can block a licence; it cannot block a power contract.
That is the thesis. The rest of this is what the thesis costs, and where it runs into things Alibaba does not control.
The model number is the least informative thing on the slide
Qwen 4 is in training now, Alibaba said. The 4.5 and 5 series are projected to scale to five trillion to 10 trillion parameters. For scale: the current flagship, Qwen3.8-Max, is a sparse mixture-of-experts model with 2.4 trillion total parameters that activates roughly 95 billion per token — about a 4 per cent activation ratio — carrying a 1-million-token context window at $2.00 per million input tokens and $6.00 per million output. Moonshot's Kimi K3 sits at 2.8 trillion. So the plan is a model two to four times larger than anything Alibaba has shipped, arriving over a multi-generation roadmap rather than a single release.
Anyone who has watched this metric for two years knows what to do with it. Total parameters count every expert stored in the checkpoint; they do not tell you how many fire per token, and they do not tell you the memory you need to serve the thing. A 2.4-trillion-parameter model at FP8 needs about 2.4 TB just for weights before a single token of context cache — roughly 18 H200-class cards, as one analysis of the July preview put it. At four bits, a 10-trillion-parameter model would want on the order of five terabytes of weights. Parameter counts are a rough gauge of stored capacity. They are not a gauge of capability, and Alibaba's own history is the proof: the open-weight Qwen3.8-27B scores 52 on the Artificial Analysis Intelligence Index, above the 47 the flagship carries.

Which brings us to the part of the keynote that is actually new.
RSI, with numbers attached
Wu said the Qwen team has made "meaningful progress" on recursive self-improvement — the loop in which a model identifies its own limits, designs experiments and synthesises data to improve itself. Labs have been vague about this for two years because it is the claim that frightens regulators; the same week that the frontier-control statement from twenty-two world leaders warned about systems circumventing testing safeguards, a Chinese cloud was publishing run counts for one.
Alibaba's numbers: over more than a month of fully automated runs covering pipeline design, data validation, iterative experimentation and error diagnosis, Qwen3.8-Max completed 33 iterative cycles, and autonomous training and post-training optimisation lifted its Artificial Analysis score from 40 to 45. In a second experiment, the model spent over 60 hours on an entire chip design lifecycle, making more than 10,000 EDA tool calls, and produced production-grade bus modules that cut chip area by 42 per cent with no performance loss.
Read that second one carefully, because it is the one that should interest people who do not care about leaderboards. Alibaba is claiming a language model shortened a step in its own accelerator programme. If it holds, that is the flywheel every chip company has been trying to close since models got good enough to write RTL — the loop where silicon trains models and models design silicon. The reference point is real progress elsewhere on the same problem: Architect Labs published a paper in August describing an accelerator designed, verified and deployed from scratch in two weeks by AI, with its own Qwen3 endpoint used to optimise the next generation of the design. The bar for first-silicon success in the industry is around 14 per cent, so any credible shortening of the loop is worth money.
Two caveats keep the claim honest. The score moved from 40 to 45 on an index whose version has changed more than once, and the numbers above it are neither the frontier nor even consistent inside Alibaba's own family — the 27B sibling scores higher. And "meaningful progress" on self-improvement is a category that currently has no verification regime, which is exactly why Ezra Klein has called for a ban on self-improving systems rather than a benchmark.
The chip is real, the memory behind it is not
T-Head's Zhenwu V900 is the concrete half of the announcement and the least arguable. It carries 216 GB of GPU memory and 1,200 GB/s of inter-chip bandwidth, supports FP8 and FP4 natively, and Alibaba says it delivers three times the performance of the M890 that shipped in May. Mass production and commercial release are scheduled for the first quarter of 2027 — a pull-in from the roadmap Alibaba gave in May, which put the V900 in the third quarter of 2027 and the J900 in 2027's third quarter again. Accelerating a chip by two quarters is not a marketing decision.
The chip lineage has a customer base, too: T-Head says the Zhenwu line has served more than 650 customers across automotive, finance, embodied intelligence, energy and manufacturing, up from the 400-plus external customers and 560,000 cumulative units it claimed at the M890 launch. That is a shipped product line, not a keynote render.
The supernode figure is where the arithmetic gets interesting. Alibaba says its upgraded supernode — Zhenwu V900 plus the ICN switch, Panmai SmartNIC and Zhenyue SSD controller — can support a cluster of up to 500,000 cards. In the same conference it launched HPN 8.0 Pro, an AI networking architecture delivering 100 petabits of bandwidth with a single cluster supporting more than 130,000 network ports at 800G. Both numbers are the company's. A 500,000-accelerator domain and a network that tops out at 130,000 ports do not obviously describe the same machine, and the keynote did not reconcile them. Every large-cluster claim of the last three years has turned on interconnect rather than silicon, which is why Nvidia's scale-up domains stop where they do; treat the half-million number as a ceiling the hardware is designed to reach, not a deployment.
Then the floor: memory. Alibaba did not say where the HBM inside a 216 GB accelerator comes from, and the honest answer from industry data is that a domestic option exists and it is not good yet. CXMT has begun pilot production of eight-layer HBM3 and is supplying samples to Chinese chip designers including T-Head and Cambricon, but front-end yields have been reported at roughly 30 per cent and stuck near 25 per cent, with only about 70 per cent of survivors qualifying after back-end processing — around 20 good stacks per 100 attempts against an industry golden yield of 80 per cent. The gap shows up in the silicon: CXMT's stacks carry roughly 3,000 through-silicon vias per die against more than 8,000 for SK hynix's HBM3, which is what "lower process complexity at the cost of bandwidth" looks like in a datasheet. SMIC, meanwhile, produces 7nm-class logic through multi-patterning on DUV with industry sources putting its 5nm and 7nm wafer prices 40 to 50 per cent above TSMC's at yields under a third.
A 500,000-card cluster, a 10-trillion-parameter model and a chip with 216 GB of on-package memory are all downstream of that. Alibaba can design the accelerator. It cannot yet buy the bandwidth in the quantity the plan assumes, and no amount of vertical integration fixes a yield curve — only production runs do.
The balance sheet is the other bottleneck
The third leg is capital, and here Alibaba has already shown its cards. In the June quarter the company spent RMB 67.68 billion (about $9.5 billion) on capital expenditure, up 75 per cent year on year, driving free cash flow to a net outflow of RMB 44.67 billion and adjusted net income down 38 per cent to RMB 20.71 billion — the profit line was the invoice for the AI buildout. In August it raised HK$80 billion, about $10.2 billion, in Hong Kong's largest-ever follow-on share placing, with 100 per cent of proceeds directed at full-stack AI; the shares fell as much as 10 per cent on the dilution. The three-year commitment stands at RMB 380 billion from February 2025, with Chinese media reporting management weighing an increase to RMB 480 billion.
Wu was candid on Tuesday about which constraint bites first: demand is "exceptionally robust," he said, but global shortages across the AI data centre supply chain are limiting how fast Alibaba can expand, and "the industry's mid-to-long-term demand far outpaces our supply capabilities." That is a supplier telling you the queue is the product. It is also the same wall the rest of the industry hit in 2026 — the AI buildout has been a fight for factory slots all year, as we laid out in the equipment queue is the new power bottleneck — and China's version of it is worse because the export controls that forced the vertical-integration strategy also cut it off from the best memory in the world.
What to watch
Three checkable things. Whether the Zhenwu V900 actually ships in the first quarter of 2027 and in what volume, since a fab-constrained launch is the whole story. Whether CXMT's HBM3 yields move off 25 per cent, which decides whether the 500,000-card cluster is a product or a target. And whether Qwen 4 arrives with a published active-parameter count and an independent Artificial Analysis score, because the 40-to-45 RSI claim is currently a company reporting on its own homework.
The strategy is not the overreach. Alibaba is right that owning the stack is the only way to be immune to the controls. The overreach is the assumption that the constraints sit where the announcement puts them — on models and chips. They sit in a memory fab in Hefei and a free cash flow line that has already crossed zero.
If self-improvement is the thing that decides this race, should companies be required to publish run counts the way they publish benchmark scores? Tell us in the comments.
Sources: Alibaba · Reuters · Bloomberg via Techmeme · The Economy — CXMT HBM3 yields · Tom's Hardware — SMIC prices and yields · Artificial Analysis — Qwen3.8 Max