September 28, 2026Updated daily by the AI editorial team
← 🤖 AI

2026-09-27

Google Rolls Out Gemini Flash TTS with 2,000+ Voices and Cloning for Generative Audio

Google has introduced two new text‑to‑speech models, Gemini Flash TTS and the lighter‑weight Flash‑Lite TTS, expanding its Gemini ecosystem into high‑end generative audio. The models are being rolled out across the Gemini API for developers, as well as tools like AI Studio, Gemini Notebook, and Google Vids, making synthesized speech a default capability alongside text and image generation.

Gemini Flash TTS ships with a library of more than 2,000 prebuilt voices and fine‑grained controls over tone, speaking rate, and emotional style. It also supports voice replication, allowing users to generate audio that mimics an existing speaker from a short reference clip. That opens attractive use cases for narration, gaming, localization, and advertising, but also raises familiar worries about impersonation and deepfake abuse. Google says it will enforce consent requirements and policy limits around sensitive uses of cloned voices.

Flash‑Lite TTS is tuned for high‑volume applications such as call centers and automated reading services, with an emphasis on low latency and cost efficiency rather than creative flexibility. As voice models become as programmable as text, the competitive race is shifting from basic speech quality to ecosystem reach, guardrails, and how convincingly platforms can claim to manage the social risks of synthetic audio.

Source: The week's top AI news | Glonce