Xiaomi's XRING O100 brings 330 tokens-per-second LLMs to devices
Xiaomi used a chip event in China on Monday to unveil the XRING O100, an accelerator built specifically for running large language models on-device — no cloud round-trip required. Announced alongside two siblings (the XRING O3 flagship processor, now in mass production, and the D100 autonomous-driving chip), the O100 is what the company calls the industry's first AI accelerator to use 6nm wafer-level 3D vertical stacking, with hybrid bonding at a 1.4-micron pitch.
The claimed payoff is bandwidth: 1.22 TB/s of memory throughput, which Xiaomi says is sixteen times what a conventional flagship phone manages, enough to push on-device LLM inference as fast as 330 tokens per second. Development of the chip is complete, with commercial products due next year; a prototype shown at the event ran it in a folding form factor with active fan cooling. Local inference is becoming the industry's standard answer to per-token bills and privacy worries, and a phone maker shipping its own accelerator raises pressure on Qualcomm and MediaTek to keep pace. One caveat worth keeping: these are vendor-claimed numbers from a launch event, so treat the 330 tokens-per-second figure as a ceiling until independent benchmarks land.
What to watch: which product ships the O100 first — Xiaomi's phone line or a dedicated edge device — and whether real-world token rates land anywhere near the demo numbers.
Would you trade a bit of bulk for a phone that runs models without calling the cloud? Tell us in the comments.
Paul Graham says that if he were 17 today he would learn to build LLMs from scratch — not start a startup. The Y Combinator co-founder's advice drew more than 600,000 views within a day, along with a pointed rebuttal from Meta's chief AI scientist: Yann LeCun replied that he would instead study why LLMs can write his essays but not clean his bedroom, then chase architectures beyond LLMs that can quickly learn physical-world skills. When two of the most-followed voices in tech draw opposite conclusions from the same moment — deep-learning craft today versus the post-LLM frontier tomorrow — that disagreement is itself the signal worth reading.
Sources: Ifeng Tech · Ifeng Tech — IT Home · Securities Times · ET CIO — Reuters · Paul Graham on X · Yann LeCun on X