Alibaba Cloud's Zhenwu M890 supernode runs 2T-parameter models
China's compute buildout just got a new centerpiece, and South Korea's flagship AI project hit an embarrassing checkpoint. Alibaba Cloud is renting out a supernode that can train trillion-scale models, while Seoul's government-backed models stumble over a two-sentence reasoning question.
Alibaba Cloud has made its Lingjun Zhenwu M890 super-node instances generally available, positioning them as the first supernode in China to run models exceeding 2 trillion parameters. The first region opened in Ulanqab, Inner Mongolia, built on T-Head's Zhenwu M890P chips with a 64-card scale-up interconnect running at 800 GB/s and a roughly 9 TB high-bandwidth memory pool. Alibaba says the platform natively handles FP8 and FP4 low-precision compute and delivers up to three times the training performance of the previous-generation Zhenwu 810E on intelligent-driving and embodied-intelligence workloads. Both Moonshot's Kimi K3 and Alibaba's own Qwen3.8 Max are already adapted and serving customers on the new instances.
The move matters because of what it offers: a fully interconnected 64-card unit on the public cloud, no hardware assembly required, aimed at large-scale MoE training and high-throughput inference for next-generation foundation models. That is the Chinese cloud answer to the supernode arms race — frontier-scale training capacity rented on demand, so labs without their own clusters can train trillion-parameter models instead of only the hyperscalers that own the iron. Watch for this becoming the default staging ground for China's next big foundation-model runs.
South Korea's "national representative" AI models are flunking a viral reasoning test — the "car wash benchmark" — as the second official evaluation enters its final stretch. Asked "The car wash is 50 meters ahead. I want to wash my car. Should I walk or drive?", multiple models from the four competing teams — LG AI Research, SK Telecom, Upstage, and startup Motif Technology — answered that walking is reasonable. The Chosun Daily's own tests on August 11 found two of the participating chatbots got the question wrong and one got it right, while Google Gemini correctly recommended driving. (The fourth team hasn't made its chatbot public.)
The stakes are real: this evaluation round adds 200 public user evaluators whose scores count for 25% of the result, and user assessments close August 12. The deeper controversy is what the test exposes about the process. Industry insiders told Chosun that some teams hired overseas training specialists to tune models against benchmark weaknesses — optimization that lifts scores without improving general reasoning — and users have posted screenshots alleging certain answers were pre-fed, "cheat sheet" style. As one insider put it, the project's goal was frontier parity, but "companies, caught in competition, seem to have shifted their primary goal to passing evaluations."
What to watch: Korea's second-round results after user evaluations close — and whether Alibaba's M890 becomes the default home for China's next trillion-parameter training runs.
If a government-backed model can't reason its way to a car wash, what is its benchmark score actually measuring? Tell us in the comments.
Sources: TMTPost · 光通信 Pro · The Chosun Daily