StepFun ships five voice models and takes two global No. 1 spots
A Chinese lab just shipped the whole voice stack in one afternoon — and the pacing fight in Washington got a very public answer from Nvidia's CEO.
StepFun released five StepAudio 3 models on the same day, and two of them now sit at the top of Artificial Analysis's speech leaderboards. StepAudio 3 Realtime, ASR, TTS, Gen and Music went live on StepFun's open platform Tuesday, covering everything from full-duplex conversation to music generation. The company claims 98.9 overall on the Artificial Analysis full-duplex conversational-dynamics ranking and 99.7% speech-reasoning accuracy, both first globally, while StepAudio 3 ASR ties for first on non-streaming transcription at a 1.7% word error rate — ahead of Microsoft's MAI-Transcribe-2 at 2.0% and ElevenLabs Scribe v2 at 2.2%. The interesting part is not the scores, it is the mechanism: a September 12 technical report from a 90-author team led by Bin Lin describes "think-while-speaking," where private reasoning runs in parallel with the audio already being spoken, so a hard question gets answered while the model is still talking instead of after a dead-air pause. StepFun's own numbers carry the usual caveat — the full-duplex latency is 8.83 seconds to first audio against 0.44 for Deepslate Opal, and the synthesis-side claims rest mostly on its own benchmarks until third-party votes land. But the direction is unmistakable: voice is following text's trajectory last year, from "can hear and speak" to "can chat and finish a task," and the names at the top of the board are Chinese.
Jensen Huang put the President on speakerphone onstage and agreed to nothing about pacing. Nvidia's CEO took a surprise call from Donald Trump during his interview at the All-In Summit in Los Angeles on Monday, where Trump declared AI takeover fears "a hoax," said "the robots will not be taking over," and argued that slowing down plays into China's hands; Huang did not push back, telling him "we're not going to let that happen" and promising to "make sure that everybody wins in the AI race in America." The value of the moment is what it says about the other side of the argument: a lab chief asking for evaluators and a pacing framework is worth nothing if the company selling the compute publicly refuses the premise. We covered the private half of this week's White House manoeuvring this morning — Trump met Altman privately at the GOP convention, then attacked Amodei.
What to watch: whether StepFun's Realtime claim survives third-party voting on the English TTS arena, where Cartesia's Sonic 3.6 still holds the top Elo.
Voice models keep landing on leaderboards rather than in products — has a speech model actually replaced your typing yet? Tell us in the comments.
Sources: StepFun documentation — StepAudio 3 ASR · arXiv — StepAudio 3 Realtime Technical Report · TMTPOST · AIBase · FREEAI.help · The Verge — Huang puts Trump on speakerphone · CNBC · Techmeme