Bilibili opens a 100-model arena — GPT-6 Astra takes the top slot
Two stories where the ranking comes from outside the lab: China's biggest video platform just turned creator reviews into a live model leaderboard, and a clinical voice company published the first benchmark of whether AI can actually say a drug's name.
Bilibili's new AI Infinite Arena ranks models by how often they win — not by how well they score. The platform opened a leaderboard fed by its UP hosts (creators), who test more than a hundred models on their own terms: code, reasoning, collaboration, knowledge, and whatever stunt they invent, with the standings updating as new reviews land. In the first round GPT-6 Astra took the most top finishes, GLM-5.3 placed second, and domestic Chinese models held three of the top five spots. Bilibili says the format is deliberately unlike a scoreboard — no fixed dimensions, no set prompts — so a model's rank reflects how often working creators reached for it and preferred it on a real workflow.
That makes it a different kind of evidence: noisier than a benchmark run, but much harder to game with benchmark-specific tuning. Bilibili has the audience to make it matter — AI knowledge watch time on the platform grew 72% year over year, with more than 190 million monthly viewers of AI content. That is now the largest consumer-facing model ranking in China, and a quiet power play: the ranking ordinary users see may not be the one labs cite. State the caveat plainly, though. The platform labels its results reference-only, and top-finish counts reward topic choice and popularity as much as raw capability. This is a taste test, not an eval.
Voice AI mispronounces up to one in three newly approved drug names. Synthio Labs, a clinical voice company, published DOSE (Drug-name Oral Synthesis Evaluation), the first public benchmark of how accurately text-to-speech systems say drug names. Nine commercial systems read the same clinical sentences covering 274 names, 146 of them recently approved; each pronunciation is scored 0–5 against verified references, with 4 and above passing. General-purpose models passed between 63.1% and 80.3% of names overall, and every one of them fell on the new ones — ElevenLabs' eleven_v3 dropped from 93.0% on established names to 67.1%, Google's Gemini TTS from 89.1% to 61.6%, and Microsoft Azure passed fewer than half of generic names. One system read Xofluza out letter by letter.
The failure mode is the story. A model can sound convincingly human and still mangle the one word in the sentence a pharmacist or patient has to act on — and the benchmark's own numbers suggest the fix is domain-specific rather than general. Synthio's RxPronounce passed 91.2% of names overall and 87.0% of newly approved ones, 10.9 points ahead of the next system, by treating pronunciation as a medical data problem instead of a language skill. The full dataset and audio are public on Hugging Face, which turns a vendor claim into a gap anyone can measure.
What to watch: whether Bilibili's creator arena starts moving model choice in China the way arena-style Elo scores move it in the West.
Would you trust a crowd-ranked leaderboard over a lab-run benchmark? Tell us in the comments.
Sources: Bilibili AI Infinite Arena · IT之家 via Sina Finance · Synthio Labs DOSE benchmark · PR Newswire · Yahoo Finance