Google's Gemini 3.8 TTS clones a voice from 30 seconds of audio

Share
Google's Gemini 3.8 TTS clones a voice from 30 seconds of audio

Google put two text-to-speech models into the Gemini API today and made voice cloning a standard developer feature. Also: Meta's Connect keynote landed three products, and Microsoft's new Gulf number is mostly an old one.

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on Wednesday, and the headline capability is replication — building a synthetic voice from a reference recording as short as ten seconds, with Google's own marketing using thirty as the round number. The models also do voice design: describe a voice in text, get a persistent voice ID back. Google says the audio family supports more than 100 languages, and its documentation is more precise than the blog — 130 for Flash TTS and 101 for Flash-Lite. Availability is broad for launch day: both models are in AI Studio and the Gemini API, Flash TTS is in Gemini Notebook, Flash-Lite is in Google Vids, and Gemini Enterprise follows later. Google claims 2,000-plus ready-made voices; Simon Willison counted 2,089 in the live catalog.

What makes this worth a second read is the consent mechanism, which is unusually concrete for a shipping feature. To clone a voice you must record the same adult speaker reading a fixed line — "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model" — and that recording has to match the reference sample. Google also says every clip carries a SynthID watermark and that replicated voices get C2PA credentials, though neither appears in the model card, so treat both as vendor claims rather than audited facts. The pricing is a cliff rather than a discount: Flash TTS runs $0.50 per million input tokens and $9.00 per million output tokens through the end of 2026, then doubles on January 1, 2027, with Flash-Lite going from $0.50/$6.00 to $1.00/$12.00. Voice replication through AI Studio is also unavailable in Illinois, Texas, the EEA, the UK, Switzerland and India — an odd patchwork that reads like a compliance map rather than a product decision.

The independent read is thinner than the launch material. The Decoder's hands-on got a convincing thick German accent out of a style prompt — and a high-pitched whine in the background in both tests, plus voice drift at the end of one clip. Google's benchmark wins on Hume AI's voice-design index are its own numbers, and the leaderboards themselves are JavaScript-gated, so we could not verify them. One clarification: this is not the realtime model we covered last week — Gemini 3.8 Live tops voice benchmarks at a sixth of GPT-Live's price was speech-to-speech dialogue; these two are text-to-speech, with different model IDs and no overlap. Our read: the consent recording is the most interesting thing here, because it asks for an affirmative act instead of a checkbox — and because a consent recording is exactly the artifact that disappears the moment a platform resells the capability.


Meta used its Connect keynote to ship a camera-free pair of AI glasses, a $1,299.99 headset that looks like glasses, and a home for its Muse agent on both. Ray-Ban Meta Audio is the first pair in Meta's line with no camera at all — 43 grams, twelve hours of battery plus forty-eight from the included case, 23 colour and lens combinations, $349, pre-ordered now and shipping October 13. The camera does not leave the lineup: Ray-Ban Meta (Gen 3) keeps a 12-megapixel camera, adds 3K video and a six-microphone array Meta says cuts more than 90% of background noise, and is available today from $449. Meta also confirmed Muse, its consumer agent, is coming to the glasses — another always-available surface for the agent it has been pushing onto phones since August, which we covered when Meta shipped Muse inside a sealed VM. The camera-free pair is the device we wrote about in September while it was still an unconfirmed plan code-named Luna; it shipped under the Ray-Ban name, and the six mics stayed.

The more consequential device is the one nobody will call glasses. Meta VR Glasses weigh about 100 grams, run a 5K micro-OLED display at 37 pixels per degree on Qualcomm's Snapdragon Reality Elite, push compute and battery to a pocket puck, and put Meta's AI agent inside the operating system so you navigate by voice, eyes and hand gestures with no controllers. They go on sale in Spring 2027 at $1,299.99, and Meta says they will be the first IMAX Enhanced certified VR device. Meta also opened four new markets for the glasses line today: Singapore, South Korea, Mexico and the UAE.


Microsoft says it will invest more than $10 billion in Saudi Arabia, Kuwait, Qatar and the UAE through 2030 — but the headline number is mostly a restatement. Bloomberg's own headline put the new money at about $2 billion and noted that the $10 billion total includes the $7.9 billion UAE pledge Microsoft made last November; The Next Web did the same subtraction. Microsoft's blog announcing the framework does not break the figure out at all. What is genuinely new is the four-country scope, a $400 million commitment to subsea and terrestrial connectivity across the Middle East by 2030, and cybersecurity and sovereign-cloud work with Gulf authorities. No capacity figure — megawatts or otherwise — was disclosed, the per-country split was withheld for security reasons, and Brad Smith said no new capital investment is planned in G42, HUMAIN or QAI. The timing is the tell: this is the first significant regional update since the Iran war began in late February, and the earlier attacks on AWS facilities in Bahrain and the UAE are why "digital resilience" is in the headline.

What to watch: whether anyone publishes an independent quality comparison of the new TTS models against the incumbent voice APIs, and whether Google's consent recording survives being wrapped by a third-party platform.

If a thirty-second sample and a spoken line are the whole safeguard, what happens when the platform reselling your voice never asks for the line? Tell us in the comments.

Sources: Google — Gemini 3.8 Flash TTS and Flash-Lite TTS · Google — Gemini API voice replication docs · The Decoder — Google's new Flash TTS models let you design AI voices from scratch · Simon Willison — Gemini 3.8 TTS Playground · Meta — Introducing Ray-Ban Meta Audio and more AI glasses styles · Meta — Introducing Meta VR Glasses · The Verge — Meta Connect 2026: the biggest news and announcements · Microsoft — Strengthening our commitment to the Middle East · Bloomberg — Microsoft pledges to spend additional $2B in Gulf region · Reuters — Microsoft plans $10 billion-plus Gulf investment · The Next Web — Microsoft's $10bn Gulf pledge