Gemini 3.8 Live tops voice benchmarks at a sixth of GPT-Live's price
Two launches Tuesday, one pattern: the fight has moved off raw model quality and onto the price of a unit of work — an hour of conversation, or a hairdresser's portfolio post.
Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on Tuesday, two voice-first models the company says top the realtime benchmarks while undercutting the closest rival voice models on cost. Both roll out starting September 15 through the Gemini API, Google AI Studio and Search Live, with enterprise access in private preview in Gemini Enterprise. They process visual input in near real time, detect and switch between 97 languages mid-conversation, and keep executing tool calls and API requests in the background while the conversation continues — Google's demos show chess play-by-play, live onboarding help and troubleshooting inside Search Live. Extended Thinking is the interesting half: it reasons and speaks at the same time, using early verbal cues to acknowledge a request and narrating progress through multi-step background tasks instead of going silent. It lands in Gemini Live, Docs, Gmail and Keep for Google AI Pro and Ultra subscribers; 3.8 Live is available to everyone in Search Live.
The quality gap is thin; the price gap is not. On Artificial Analysis' Speech-to-Speech Quality Index, Google puts 3.8 Live Extended Thinking at 82.6%, ahead of GPT-Live-1 Astra (Medium) at 81.5% and Grok Voice Think Fast 2.0 (High) at 81.3%. On the agentic τ-Voice benchmark it claims 68.6% against 67.9% and 56.5%, and on Sierra's τ³-Banking leaderboard 35.1% against 32.0% and 16.5%. Then the pricing: Google's cost-per-hour-of-input-audio figure is $0.84 an hour for standard 3.8 Live and $3.50 for the higher-effort Extended Thinking model, versus $4.80 for Grok Voice Think Fast 2.0 and $5.83 for GPT-Live-1 Astra. Standard 3.8 Live scores 76.0% on the same quality index — roughly a sixth of the GPT-Live price for most of the conversational ability. All audio output carries Google's SynthID watermark.
Why it matters: a one-point benchmark lead is inside the noise, and Google knows it — the pitch is that an hour of competent voice costs a sixth of the incumbent's rate, which is how you commoditize a layer before a winner locks in. The τ³-Banking number is the reality check: the best voice agent in the field resolves only about a third of banking-style tasks. This is also the third Gemini release in two weeks, after 3.8 Flash — Google ships Gemini 3.8 Flash, gating its cyber model behind Fairwind — so the model cadence is now a pricing strategy, not a research milestone.
Meituan used its Top 100 Hairdressers conference in China on Tuesday to launch 手艺人Agent ("Artisan Agent"), its first AI assistant built for individual service workers rather than consumers or enterprises. The agent handles online operations Q&A, business-data analysis, reminders and portfolio management; publishing a set of work photos used to take 11 separate steps — choosing images, writing copy, filling in service items and styles, picking tags — and now takes two actions, with the agent identifying the photos, drafting copy in the stylist's own style and pre-filling the rest for a one-tap confirmation. Meituan says daily active artisans on the platform are up more than 16% year on year, and the share of active hairdressers posting work monthly rose from 21.9% to 39.8%.
Why it matters: this is the agent deployed at the supply side of a marketplace — the listing density problem, solved by automating the chore that kept stylists from publishing. It is also a template: if a one-person service business can get an operator that answers "what should I post, and how," the same pattern ports to restaurants, repair shops and every other category Meituan already indexes.
What to watch: whether OpenAI, Anthropic and xAI answer Google's voice pricing within the week, and whether Meituan's artisan-agent pattern shows up in other service categories.
If an agent could run the online side of your business — posting, pricing, replying — would you hand it the keys, or would you rather keep the manual steps? Tell us in the comments.
Sources: Google DeepMind · Gemini 3.8 Audio model card · Unite.AI · OfficeChai · Gemini Live API docs · 雷峰网 Leiphone · Gelonghui via Futu News · 10jqka News