China's biggest edge-AI unicorn ModelBest files for A-share IPO

Share
China's biggest edge-AI unicorn ModelBest files for A-share IPO

The AI story this cycle is usually written in data centers — but two items overnight point elsewhere: China's largest on-device model maker is heading for the public market, and a new paper hands diffusion language models scaling rules of their own.

Beijing ModelBest, China's biggest edge-AI unicorn, has officially started the A-share IPO process, signing a tutoring agreement with CITIC Securities that was posted to the securities regulator's website on August 12. The Tsinghua University NLP lab spinout, founded in 2022 and known for its MiniCPM family, has raised more than 5 billion yuan this year and now carries a valuation above 20 billion yuan (about $2.8 billion) — the largest publicly disclosed valuation in China's on-device AI sector. Its backers read like a roll call of state capital: China Telecom, Shenzhen Capital Group, China Reform Holdings and CITIC affiliates. The commercial numbers behind the filing are the real signal: 38 million open-source downloads, its models shipping in more than 300,000 mass-produced vehicles, and a spot as the first domestic edge model inside an international phone maker's supply chain. CEO Li Dahai's line from this year's WAIC — "2026 is the year edge AI scales" — now has a stock listing attached to it. The IPO matters beyond one company: edge inference is the philosophical counterweight to the trillion-dollar data-center buildout, and ModelBest's path to public markets gives the on-device thesis its first liquid benchmark.


A Renmin University–Ant Group team has published the first large-scale scaling study for mixture-of-experts diffusion language models, training LLaDA MoE v2 — a 30-billion-parameter model with 3 billion active — from scratch on 23.5 trillion tokens. The paper, posted to arXiv this month, reports that the model approaches Qwen3's 30B-A3B on knowledge, reasoning and coding benchmarks while using roughly 65 percent of Qwen3's pretraining tokens, and that after supervised fine-tuning alone it beats SDAR Chat on seven of eight reasoning and coding benchmarks. The contribution that matters more than the scores is methodological: rather than borrowing autoregressive scaling recipes wholesale, the authors derive MoE-specific rules — the optimal batch size grows faster, and compute allocation diverges from what holds for autoregressive models. Diffusion language models, which generate text by denoising rather than predicting token-by-token, have spent two years as an interesting demo stuck at small scale; this is the strongest evidence yet that they can compete at frontier scale, and that the recipe is not a copy-paste of the autoregressive playbook.

What to watch: whether other edge-AI players follow ModelBest to market — and whether diffusion language models graduate from research curiosity to production workhorses.

Edge models are going public while diffusion models chase the frontier — which bet do you like? Tell us in the comments.

Sources: NetEase (163.com) · Sohu — ModelBest IPO · arXiv — LLaDA MoE v2 · Hugging Face Papers · AI Weekly · Sohu — LLaDA MoE v2 · Turing Post