Google released the models Gemini 3.8 Flash TTS and Flash-Lite TTS with voice creation from text descriptions and voice cloning
Google has published details about the models Gemini 3.8 Flash TTS and Flash-Lite TTS (GA from 22. 9. 2026): over 100 languages, voice creation from a text description, voice cloning from a 30-second recording, 2000+ preset voices, SynthID watermark.
Google has clarified the capabilities of the newly released text-to-speech models Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Both models support more than 100 languages. Flash TTS can create an entirely new voice based on a text description of the role, accent and vocal characteristics; those who do not want to design a voice from scratch can choose from a library of over 2000 preset voices, including regional variants (for example, Mexican Spanish, Quebec French and Scottish English). The voice cloning feature creates a voice profile from a 30-second audio recording, and according to Google, a recording of spoken consent from the person matching the voice sample must also be uploaded. According to Google, every generated audio clip carries an inaudible SynthID watermark for detecting AI-generated speech. Google also announced a “Voice Remixing" feature for adjusting the timbre, pitch, tempo and accent of library voices, but it is not yet available.
The models became generally available (GA) on 22 September 2026, alongside the publicly accessible Voices endpoint (/v1beta/voices) in Gemini API. Gemini 3.8 Flash TTS is intended for creative uses such as podcasts, audiobooks or game characters and, according to Google, aims for studio quality with nuanced acting and regional dialects. Gemini 3.8 Flash-Lite TTS replaces the older model gemini-3.1-flash-tts-preview and is optimized for lower-cost speech generation at scale — dubbing, audio content and real-time voice agents. Both models support directing instructions for individual lines, dialogue with two voices in one script and nonverbal sounds such as laughter or sighs.
The models are being made available through Gemini API and Google AI Studio, with Flash TTS also available in Gemini Notebook and Flash-Lite TTS in Google Vids; access via Gemini Enterprise API is to follow. According to Google, data from the free tier is used to improve products, while data from the paid tier is not. Paid pricing is listed in USD per million tokens: for Flash TTS, text input is 0.50 USD (until the end of 2026, then 1.00 USD from 1. 1. 2027) and audio output is 9.00 USD (then 18.00 USD); for Flash-Lite TTS, text input is 0.50 USD (then 1.00 USD) and audio output is 6.00 USD (then 12.00 USD). According to Google, one second of generated audio corresponds to 25 audio tokens, so an hour of audio costs approximately 0.81 USD for Flash TTS and 0.54 USD for Flash-Lite TTS until the end of 2026, then 1.62 USD and 1.08 USD from 2027 (text input is billed separately).
Why it matters
Content creators and developers can design a custom voice using just a text description or clone a specific voice from a short recording, without needing a professional recording studio or an actor recording each language separately. For companies deploying dubbing, audio content or voice agents at scale, the published price per hour of audio and the fact that it will double from January 2027 are significant — this affects cost calculations for projects planned over the long term.
Release card
Gemini 3.8 Flash TTS
- Inputs
- text input, audio (speech) output
- Price
- Text input 0.50 USD per million tokens (until the end of 2026, then 1.00 USD from 1. 1. 2027); audio output 9.00 USD per million tokens (then 18.00 USD from 1. 1. 2027)
- Availability
- Gemini API and Google AI Studio (GA from 22. 9. 2026), also Gemini Notebook; access via Gemini Enterprise API is to follow, according to Google
- podcasts
- audiobooks
- game characters
- The Voice Remixing feature for adjusting the timbre, pitch, tempo and accent of library voices is not yet available.
- In a test by an editor at The Decoder, a stylized accent produced high-frequency background noise, and in one clip the voice changed at the end.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
What was added since the original report
Verified updates
-
The models support more than 100 languages; Ability to generate voices from text descriptions; Voice cloning from 30-second audio samples; Over 2000 preset voices with regional variants; Integrated SynthID watermark for detecting AI speech
- The models support more than 100 languages
- Ability to generate voices from text descriptions
- Voice cloning from 30-second audio samples
- Over 2000 preset voices with regional variants
- Integrated SynthID watermark for detecting AI speech
Two audiences, two different impacts
What this means
For individuals
Creators of podcasts, audiobooks or game dialogue can try designing a custom voice through a text description or cloning a voice from a short recording in Gemini API and Google AI Studio.
For a business
Companies planning dubbing, audio content or voice agents at scale now have clear pricing for text and audio output, including the price increase from January 2027, which affects operating cost estimates.
DevelopmentCheck the original
Event sources
independently confirmed · 2 publishers, 1 independent. We count feeds from the same owner only once.