Gemini 2.5 models will shut down in October 2026. To avoid service disruptions,
update to a newer model (like
gemini-3.7-flash or gemini-3.1-flash-image).
Any stable Gemini Live API 2.5 models are not impacted.
Learn more.
Both text-to-speech (TTS) models and Live API models are low-latency,
speech-generating models that can be configured for different response voices
and languages. However, they serve very different use cases.
Text-to-speech (TTS) generation is a
unidirectional, request-response interaction (text in, audio out). It's
tailored for scenarios that require exact recitation of the provided
text with fine-grained control over style and sound, such as podcast
narration, audiobooks, or reading articles aloud.
Live API generation supports
bidirectional streaming for real-time voice conversations (voice in,
voice out). It excels in dynamic conversational contexts where the model
decides the applicable speech to return. Note that the latest Live API
models also support video and image input.
[[["Easy to understand","easyToUnderstand","thumb-up"],["Solved my problem","solvedMyProblem","thumb-up"],["Other","otherUp","thumb-up"]],[["Missing the information I need","missingTheInformationINeed","thumb-down"],["Too complicated / too many steps","tooComplicatedTooManySteps","thumb-down"],["Out of date","outOfDate","thumb-down"],["Samples / code issue","samplesCodeIssue","thumb-down"],["Other","otherDown","thumb-down"]],["Last updated 2026-08-19 UTC."],[],[]]