Physix Frontier · News Briefing Card (IT Home · Sep 23, 2026)

Google launches Gemini 3.8 Flash text-to-speech models

KEY FACTS

  • Google introduced two text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS.
  • Flash TTS targets character design and creative scenarios, while Flash-Lite targets high-volume, low-cost voiceover.
  • The models support more than 100 languages, and the voice library expands from 30 to more than 2,000 ready-made voices.
  • Voice cloning requires only a 30-second audio sample, with consent verification and SynthID watermarking added.
  • Both models have rolled out to developers in the Gemini API and Google AI Studio.

KEY DATA

71.4 分Hume AI voice design benchmark total score
60.8 分Hume AI accent modeling score
100 多种Number of supported languages and dialects
2000 多种Size of ready-made voice library

PHYSIX OBSERVATION

Speech synthesis is shifting from picking preset timbres to finely orchestrating lines, which is a real productivity tool for audiobooks, podcasts, and game voiceover. But with the voice cloning threshold dropping to 30 seconds, plus regional ban lists, it shows compliance pressure has already preceded product rollout; what the industry competes on next is not just audio quality, but authorization and traceability mechanisms.

Source: IT Home report