AI Models & Platforms

Google Rolls Out Gemini 3.8 Speech Models In API And AI Studio

mm
Add Unite.AI to your preferred sources on Google

Google on September 23, 2026 introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two text-to-speech models it described as its most expressive audio generation models yet. Both models began rolling out the same day in the Gemini API and Google AI Studio.

The announcement, written by Group Product Manager Leland Rechis and Director of Research Science Alan Cowen on behalf of the Gemini Audio Team, positions Gemini 3.8 Flash TTS for deep creative direction and character design across gaming, immersive audiobooks, podcasts, and interactive media. Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient use, optimized for dubbing, audio content creation, and voice agents with fine-grained control over tone, pacing, and expressive nuance. The announcement presents the pair as following earlier Gemini Audio releases: 3.5 Live Translate, 3.5 Transcribe, and the Gemini 3.8 Live and 3.8 Live Extended Thinking models. According to the Gemini 3.8 Audio model card, the models are based on Gemini 3 Pro, and the TTS variants accept text input up to 8K tokens and return audio output up to 64K tokens.

Voice Design and Replication

Google said Gemini 3.8 Flash TTS can create bespoke voices from scratch through natural language prompting, customizing role, accent, and voice characteristics across more than 100 languages and dialects. The company said the release scales its original set of 30 voices into a library of more than 2,000 production-ready voices with broad language coverage, including regional varieties such as Mexican Spanish, Quebec French, and Scots English. Users can save and manage their custom voices to keep performance consistent across ongoing projects, and a voice-remixing feature that adjusts timbre, pitch, pace, and accent on library voices is listed as coming soon.

Voice replication recreates a consistent vocal profile from a 30-second audio sample of the user’s own voice or a voice the user has rights to use. Google said the capability ships with built-in consent verification, SynthID watermarking, and C2PA credentials.

Performance Direction and Reported Benchmarks

Both models allow line-by-line direction of each performance: users can write their own stage directions or let Gemini steer delivery from natural script cues. Google said long-form generation holds voice quality, pacing, and character timbre steady across hours of continuous audio with minimal speaker drift, aimed at podcasts and audiobooks. Native two-speaker scene staging directs multi-turn conversations from a single script while keeping the two voices separated with natural turn-taking. Scripted vocal bursts and backchanneling add non-verbal cues such as <laughs>, <sigh>, and <gasp>, plus interjections like “mhm” and “yeah.”

On evaluations, Google reported that Gemini 3.8 Flash TTS took the top overall position on Hume AI’s Voice Design Benchmark with a score of 71.4 and also led accent modeling at 60.8. It said Flash TTS and Flash-Lite TTS hold the first and second positions on Hume AI’s Overall Quality Index, and that the new models show major improvements over Gemini 3.1 Flash TTS across use cases including long-form content and dual-speaker screenplay control. In blind human preference evaluations on Voice Arena, Google said the two models took top positions among competitors in languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.

Safety, Limitations, and Availability

Under the consent verification system, a user must provide a verbal consent recording from the voice owner that matches the reference speaker before a replicated voice can be created. Every audio clip the Gemini Audio models generate carries a SynthID watermark, which Google describes as imperceptible and embedded directly in the audio output so that AI-generated speech remains detectable.

The model card lists known limitations including hallucinations and occasional slowness or timeout issues, and it gives a knowledge cutoff of January 2025. For frontier safety, Google DeepMind said its assessments found no meaningful new capabilities or material performance increases in the Gemini 3.8 Audio models compared with Gemini 3.7 Flash, which its evaluations found did not reach any Tracked or Critical Capability Levels under its Frontier Safety Framework. On that basis, it said the Gemini 3.8 Audio models are not likely to reach those levels.

Gemini 3.8 Flash TTS is also available in Gemini Notebook and Flash-Lite TTS in Google Vids, with enterprise access via API in Gemini Enterprise listed as coming soon. Google AI Studio adds an audio playground built like a voice design workspace, where developers can prompt new vocal identities from scratch or replicate a voice and then direct line-by-line delivery in a dual-speaker screenplay editor. A footnote in the announcement notes that voice replication through AI Studio is not available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland, and India. Developer platforms Agora, LiveKit, Pipecat, and Vercel support the models through the Gemini API, and Google named Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang as partners integrating the new TTS models for global dubbing, media localization, and conversational voice agents.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.