Google is expanding Gemini beyond text and images with two new text‑to‑speech models that feel less like traditional TTS and more like a small audio studio inside the API. Announced on September 23 and rolling out through the Gemini API and Google AI Studio, Gemini 3.8 Flash TTS and Gemini 3.8 Flash‑Lite TTS are designed to cover the full spectrum of use cases: from cinematic narration and character voices to high‑volume dubbing and real‑time voice agents.
The pitch is simple. If you want maximum expressiveness and fine‑grained creative control, you use Flash TTS. If you need to generate hours of audio at scale with low latency and lower cost, you use Flash‑Lite TTS. Both support more than 100 languages and dialects and draw from a library of over 2,000 production‑ready voices, but they’re built for different jobs in the pipeline.
Gemini 3.8 Flash TTS: design voices from scratch
Gemini 3.8 Flash TTS is positioned as Google’s flagship creative TTS model. It’s engineered for studio‑grade voice fidelity, nuanced acting, regional dialects and long‑form multi‑speaker stability. Think audiobooks, game dialogue, podcast narration, interactive media and anything where the voice itself is part of the product.

The standout feature is generative voice design. Instead of picking from a fixed list, you can describe the voice you want in natural language: role, accent, age impression, tone, pacing, even character archetype. The system then generates a persistent custom voice persona with a preview sample and a reusable voice ID. You can refine it over time, save it to a project and call it across sessions.
Flash TTS also supports:
- Voice replication from a 30‑second sample, with explicit consent controls. Upload a short recording of an authorized speaker and the model can reproduce that voice for scripts, dubbing or read‑aloud features.
- Multi‑speaker dialogue and performance cues. You can direct different lines to different voices, specify emotions, pacing, laughter, sighs and other non‑verbal cues, and maintain consistency over hours of audio.
- Fine‑grained control over timbre, pitch, pace and accent, including a coming “voice remixing” feature that lets you start from a library voice and modify its characteristics via prompts.
For creators and studios, this turns TTS from a utility into a creative tool: you’re not just converting text to speech, you’re directing performances.
Gemini 3.8 Flash‑Lite TTS: built for volume and real‑time agents
Not every use case needs cinematic nuance. Some need throughput, latency and cost efficiency at scale. That’s where Gemini 3.8 Flash‑Lite TTS comes in. Google describes it as the replacement for the earlier 3.1 Flash TTS preview, optimized for high‑volume production, conversational voice agent cascades, read‑aloud features and everyday single‑speaker generation.
Flash‑Lite TTS is designed for:
- High‑volume dubbing and content libraries, where you might be generating thousands of hours of audio across many languages and need predictable quality and pricing.
- Real‑time voice agents, such as customer support bots, interactive assistants and in‑app narration, where latency and cost per request matter more than theatrical range.
- Reliable voice replication for approved speakers, enabling personalized read‑aloud experiences without the full creative overhead of Flash TTS.
In practice, many teams will use both: Flash TTS for hero content (trailers, key scenes, signature brand voices) and Flash‑Lite TTS for the long tail of everyday audio.
How developers and enterprises can deploy these models
Both models are generally available now through the Gemini API and Google AI Studio, with a dedicated Voices endpoint for listing, creating and managing custom voices.
For developers:
- You can call the TTS models directly from the Gemini API, specifying either a prebuilt voice ID or a custom voice you’ve created via prompt.
- The API returns high‑fidelity audio files suitable for direct use in apps, games, podcasts and video workflows.
- Voice design is accessible in Google AI Studio, where you can experiment with prompted voices, save them and then reference them in your code.
For enterprises:
- Gemini 3.8 Flash TTS is rolling out first to developers and AI Studio users, with broader availability coming soon via API in Gemini Enterprise for governed, organization‑wide deployments.
- Flash‑Lite TTS is already available alongside tools like Google Vids, positioning it for internal comms, training content and large‑scale corporate audio needs.
Google emphasizes that all audio generated by these models is watermarked with SynthID, its inaudible watermarking technology, to help identify AI‑generated speech and support transparency.
Where this fits in the broader AI audio race
Google isn’t the only player pushing into advanced TTS, but the combination of custom voice design, voice cloning with consent, multi‑speaker direction and a clear split between creative and efficiency‑focused models sets a new baseline for what’s expected from a platform TTS service.
For content creators, this means you can build a library of brand‑specific voices and character personas without hiring a full voice cast for every project. For product teams, it means voice agents can sound less robotic and more like natural conversational partners, with distinct personalities and consistent delivery over long sessions.
The real test will be how these models perform in production at scale: latency, cost, reliability and how well the custom voices hold up across languages and use cases. But with Flash TTS and Flash‑Lite TTS now generally available, Google is clearly betting that audio will be a core part of the next generation of AI applications and that developers will want the same kind of creative control over voices that they already have over text and images.
Mohit Sharma
Mohit Sharma is the Founder of AISEOToolshub and an SEO & Digital Marketing Expert with over 6 years of experience helping websites improve their search visibility and organic growth. Mohit closely follows the latest developments in artificial intelligence and regularly shares practical insights on new AI tools, industry updates, and breaking AI news.
Discover more from AI News Hub
Subscribe to get the latest posts sent to your email.