vivo's 3B on-device model tops phone LLM chart, nears cloud scores

Share
vivo's 3B on-device model tops phone LLM chart, nears cloud scores

A SuperCLUE benchmark released Friday found that a small model living entirely on a phone is competitive with the cloud. vivo's BlueLM 3.5 Nano 3B took the top spot on the group's on-device leaderboard with 89.86 points, edging past the field while scoring within a small margin of flagship cloud systems it wasn't even competing against.


vivo's on-device BlueLM 3B topped SuperCLUE's phone-LLM chart — and nearly matched the big cloud models. The 3.5 Nano scored 89.86 in SuperCLUE's OnDevice run, the highest among phone-local models. Cloud models were listed only as reference points and didn't rank: Google's Gemini 3.6 Flash hit 93.64, ByteDance's Doubao Seed 2.1 Pro scored 92.99, and Alibaba's Qwen3.8 Max landed at 92.25. Behind vivo on the device table sat Qwen3.5 9B at 87.82 and Qwen3.5 4B at 85.16, with GLM 4.6V Flash, Gemma 4 E4B, and Gemma 4 E2B further down.

What makes the result notable is how small the winning model is. On-device models have long been written off as a significant step below cloud quality, held back by phone hardware limits on parameter count. A single 3B model coming this close to a frontier cloud model — roughly a 4-point gap to Gemini 3.6 Flash — suggests the "you need a huge model" assumption is weaker than it looked. The tradeoff is real, though: a 3B model still hits a ceiling on the hardest tasks, and one benchmark run isn't proof of everyday performance across apps and system tuning. But the privacy win is structural — everything runs and stays on the handset, with no upload, which is exactly where on-device AI is heading as images and text-heavy assistants eat into the phone's weekly compute budget.

What to watch

Whether the gap narrows further as phone vendors race to optimize for these leaderboards — and whether a genuinely offline assistant becomes a default phone feature, not a Telephony-flavored extra.

Does a 90-point on-device model make you want your next AI assistant purely offline, or do you still trust the cloud more? Tell us in the comments.

Read more

Lambda raises up to $4B from Blackstone ahead of its IPO

Lambda raises up to $4B from Blackstone ahead of its IPO

The neocloud money is consolidating fast, and today's inbox shows both ends of the market: a heavyweight pre-IPO round on one side, and a Google open model you can run on a phone on the other. Lambda is raising up to $4 billion led by Blackstone and Coatue at a $14.5 billion pre-money valuation — its last private round before a planned IPO. The Wall Street Journal reported the scoop from a letter to limited partners, and Reuters independently confirmed the headline terms: the round is led by t

South Korea bets $3.49B on its own frontier AI model

South Korea bets $3.49B on its own frontier AI model

Sovereign-model money is getting serious, and the hardware money is following it. Today's inbox: Korea's nine-figure upgrade to its homegrown model push, a physics-simulation startup priced like a chip designer, and Google turning a geospatial model loose on public health. South Korea is putting 4.7 trillion won — about $3.49 billion — of state equity behind a homegrown frontier AI model. The Ministry of Science and ICT confirmed the figure as part of its proposed 2027 budget, split into two t

Mistral's Le Chonk puts Europe's sovereignty bet on a download date

Mistral's Le Chonk puts Europe's sovereignty bet on a download date

Mistral's biggest model ever is real, benchmarked and for sale today — but the thing that makes it matter to Europe's sovereignty argument, the weights, is still three weeks out. The preview settles who built it; the release will settle whether it counts. What Mistral actually shipped Mistral opened a public preview of Mistral Large 4 — unofficially ML4, officially le Chonk — a 1 trillion-parameter mixture-of-experts model with 49 billion active parameters and native multimodal input. The p

Mistral unveils Le Chonk: a 1T-parameter open-weights model

Mistral unveils Le Chonk: a 1T-parameter open-weights model

The biggest open-weight release outside China lands in public preview today, and the country that spent the week promising its own frontier model just put a price on the ambition. Mistral has opened a public preview of Mistral Large 4 — codenamed "le Chonk" — a 1 trillion-parameter mixture-of-experts model with 49 billion active parameters, natively multimodal, which the company calls its largest and most capable model to date. The preview API is live today on Mistral Studio at $1.36 per milli