AI RACE— The AI Race
New Models

Google Launches Gemini 3.8 Flash TTS Models with Text-Prompted Voice Design

Google has introduced Gemini 3.8 Flash TTS and Flash-Lite TTS, two new speech generation models that support over 100 languages and allow creators to generate custom voices using text prompts or short audio samples.

09/24/2026, 00:39
Google ra mắt Gemini 3.8 Flash TTS: Tạo và tùy biến giọng nói AI từ câu lệnh văn bản

Google Unveils Gemini 3.8 Flash TTS and Flash-Lite TTS

Google has expanded its generative audio portfolio with the release of Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The two text-to-speech models cover more than 100 languages and introduce granular controls for vocal delivery, including line-by-line stage directions.

Gemini 3.8 Flash TTS is positioned for expressive creative workflows—such as video game character dialogue, podcasts, and audiobooks—and allows developers to generate entirely new synthetic voices from written descriptions. Gemini 3.8 Flash-Lite TTS is tailored for scalable, budget-conscious deployments, including localization dubbing, automated voice agents, and high-volume media production. Both models are rolling out through Google AI Studio and the Gemini API, with enterprise access scheduled to follow.

Text-to-Voice Design, Script Control, and Pricing Breakdown

Flash TTS introduces the ability to describe vocal attributes—including pitch, accent, tone, and character persona—directly via text prompts. Users who prefer pre-existing options can select from a catalog of over 2,000 preset voices spanning regional accents like Mexican Spanish, Scottish English, and Quebec French. An upcoming feature dubbed "Voice Remixing" will allow creators to tweak library voices by modifying timbre, tempo, and pitch.

For voice replication, Flash TTS includes a cloning tool that builds a custom profile from a 30-second audio recording. To prevent unauthorized impersonation, the system requires the speaker to record an explicit verbal consent statement that must match the biometric profile of the source sample.

Both models support detailed performance direction:

  • Scripted Delivery: Prompts can interpret narrative cues automatically or execute explicit line-by-line stage directions, inserting pauses and nonverbal cues like laughter, sighs, and hesitations like "mhm."
  • Multi-Voice Dialogue: A dual-speaker mode enables two distinct voices to alternate within a single script. Google states the models maintain vocal consistency across extended generation sessions with minimal "speaker drift."
  • Tool Availability: Flash TTS is integrated into Gemini Notebook, while Flash-Lite TTS is accessible within Google Vids. Several developer platforms, including LiveKit, Agora, Pipecat, and Vercel, have launched support via the Gemini API.

Google structures its API pricing around text input tokens and audio output tokens, measuring audio generation at 25 tokens per second (or 90,000 tokens per hour). Through the end of 2026, text input for both models is billed at $0.50 per million tokens. Audio output costs $9.00 per million tokens for Flash TTS (approximately $0.81 per hour of generated speech) and $6.00 per million tokens for Flash-Lite TTS ($0.54 per hour). On January 1, 2027, rates are scheduled to double to $1.00 per million text tokens, with audio outputs rising to $18.00 per million tokens for Flash TTS ($1.62 per hour) and $12.00 per million tokens for Flash-Lite TTS ($1.08 per hour). Data submitted under the paid API tier is exempt from product training, whereas free-tier data may be utilized by Google.

Watermarking Standards and the Synthetic Voice Market

To address safety and provenance concerns surrounding synthetic speech, all audio generated by the new Gemini TTS engines includes Google's imperceptible SynthID watermark, designed to verify machine-generated audio downstream.

Google claims the new Gemini 3.8 TTS lineup leads across most evaluation categories on benchmarks established by Hume AI. The release places Google in direct competition with specialized voice synthesis providers like ElevenLabs and Cartesia, targeting developers building real-time conversational agents, interactive gaming assets, and automated dubbing pipelines.

◗ Sources

The Decoder09/24

Related stories