Same Qwen3 voice model, 14x faster — hosting is the product now
A YC-backed inference shop just posted numbers showing the model on the label tells you almost nothing about how it performs. Meanwhile, a music startup raised money on the opposite bet: that abundance of AI-generated audio makes human taste worth more.
Nari Labs reported the top scores on Coval's voice AI leaderboards, and the sharpest number in the release isn't its own — it's what happens to the model it didn't tune. The YC-backed host serves a 1.7B Qwen3-TTS endpoint at 8.8% word error rate and 692 ms median time-to-first-audio on Alibaba's official Flash Realtime endpoint; the same weights on Nari's stack land at roughly 50 ms first audio. Baseten's dedicated Qwen3-TTS endpoint sits in between at 6.0% WER and 101 ms. In Nari's own concurrency benchmarks, tuning tricks like trimming the silent lead-in and ramping chunk sizes brought open serving engines down to the 50 ms range at one request per second — but they all degraded past 100 ms by 6 RPS, while Nari's scheduler held sub-50 ms through 10 requests per second on a single H100, at a stated compute cost around $2 per million characters.
The mechanism matters more than the leaderboard: Nari splits the model's three stages so urgent work can cut in front of a batch, and prioritizes the first audio packet above everything else. Which means two companies advertising "Qwen3-TTS" are not selling the same product, and the model card is the least informative thing on their pricing page. Only Fluxions' 300M-parameter vui clocked a faster median first audio — a smaller model trading capability for speed, the other axis of the same trade. It's the serving-side version of what we saw on pricing — China's labs raise API prices while third-party hosts slash them.
A Vinyl Bar in Shibuya raised a $5.5 million pre-seed to build music apps that deliberately don't lean on AI generation. Ex-Spotify head of innovation Máuhan M Zonoozy's startup — backed by Mantis VC, SV Angel, Boxgroup and Dawn Ostroff, who joined as an advisor — is shipping a "musical sandbox mixer" that lets people play with the parts of music-making, including building instruments. Zonoozy's argument: streaming solved access, but interaction with music is still mostly consumption, and younger audiences already create through games and filters. He's fine with AI in the infrastructure; he's not betting the product on song generation — if AI makes generated audio abundant, taste and participation get more valuable, not less. Our take: this is contrarian while Pocket FM scales on AI-written audio, but "AI everywhere except where the user touches it" is a coherent consumer wedge, not a gimmick.
What to watch: whether Coval-style per-endpoint leaderboards become how voice APIs get bought — model name is becoming the least interesting field on the pricing page.
If two vendors serve the same open model at 14x different latency, what should you actually be buying — the weights, the stack, or the SLA? Tell us in the comments.
Sources: Nari Labs · Coval Voice AI Benchmarks · explainx.ai on Nari's Qwen3-TTS stack · TechCrunch