AI · Speech generation

Gemini 3.8 TTS in production: voices, streaming and consent

A production guide to Gemini 3.8 Flash TTS and Flash-Lite covering model routing, streaming audio, reusable voices, consent, caching and evaluation.

TOPIC HUBAI, RAG & Vector Search
Conceptual editorial illustration of two text-to-speech processing lanes converging into a secured reusable voice identity and an audio output waveform; not a Google interface.
An editorial interpretation of the topic, followed by a practical execution diagram.

Gemini 3.8 Flash TTS and Flash-Lite TTS are production text-to-speech models, not general chat models with an audio switch. Google released both to general availability on 22 September 2026 alongside the Gemini API Voices endpoint. Flash prioritizes acoustic fidelity, nuanced acting and long-form stability; Flash-Lite targets high-volume, latency-sensitive synthesis. The useful production design is therefore a routed audio pipeline with an explicit transcript contract, authorized voice identity, streaming format handling, deterministic caching and quality evaluation. It is not enough to replace a model name and listen to one demo.

Google's documentation was verified on 6 October 2026. Model capabilities, limits and voice-retention rules below are provider-documented facts. Routing policies, storage boundaries, cache keys, evaluation gates and rollout recommendations are engineering analysis.

What Google actually launched

The Gemini API release notes record `gemini-3.8-flash-tts` and `gemini-3.8-flash-lite-tts` as generally available on 22 September 2026. The same release introduced the `/v1beta/voices` endpoint for browsing voices, designing a persistent persona from text and replicating a consenting adult speaker.

Both model cards accept text and return audio. Each lists an 8,192-token input limit and a 16,384-token serving output limit, supports caching, Batch, Flex and Priority inference, and does not support function calling, the Live API, structured outputs, search grounding or thinking. Flash documents automatic detection across more than 130 languages; Flash-Lite lists more than 100. Language coverage is not the same as equivalent accent, pronunciation or expressive quality, so test the exact locale and content.

Choose TTS instead of Live audio for controlled recitation

Google distinguishes TTS from the Live API. Use Gemini 3.8 TTS when the product already has the exact words and needs controlled recitation: narration, accessibility read-aloud, localized announcements, podcast segments or the final speech stage of a voice agent. The input is text-only and the output is audio-only.

Use a Live audio model when the model must listen, reason over multimodal input and converse bidirectionally inside one realtime session. Do not send a microphone stream to a TTS endpoint, and do not choose Live solely to read known text. A common agent architecture uses a reasoning model to decide the response, a policy layer to approve the exact transcript, and TTS to speak that frozen text. This separation makes logs, redaction and retries easier to reason about.

Route Flash and Flash-Lite by measurable workload

The two Gemini 3.8 TTS models share the same request schema, so routing can be an application decision rather than two implementations. Use Flash for studio-grade narration, difficult pronunciations, regional or minority dialects, expressive multi-speaker dialogue and long-form material where voice or room-tone drift is costly. Use Flash-Lite for bulk generation, read-aloud, routine notifications and the per-turn speech stage of high-volume agents.

Do not hard-code prestige routing. Start with a workload label such as `creative_long_form`, `transactional_short`, `accessibility_read_aloud` or `agent_turn`, then verify quality, latency and cost against a representative set. A short account-balance sentence may gain nothing from the heavier model; a branded audiobook may lose coherence when optimized only for throughput. Keep a fallback that preserves the same transcript and voice authorization.

A production TTS path separates transcript preparation, model routing, voice authorization, streaming synthesis, format validation, delivery caching and measurable quality.
A production TTS path separates transcript preparation, model routing, voice authorization, streaming synthesis, format validation, delivery caching and measurable quality. Open for a larger view

Make the transcript an immutable contract

Gemini 3.8 treats the input text as a verbatim transcript. Put sustained delivery instructions such as pace, emotion or speaking style in a `speech_metadata.style` annotation. Use inline angle-bracket tags only for point events such as `<laugh>`, `<sigh>` or `<short pause>`. Google recommends English tag names even when the spoken transcript is not English.

Freeze the transcript before synthesis and store a content hash. Redact secrets, expand ambiguous abbreviations, normalize numbers and decide how the locale should pronounce dates, currency and product codes. Never let retrieved documents insert hidden style instructions or a different voice ID. The model should speak approved content, not reinterpret untrusted context.

{
  "transcript_hash": "sha256:...",
  "locale": "ar-XA",
  "model": "gemini-3.8-flash-lite-tts",
  "voice_ref": "voice_catalog:v7",
  "style_version": "support-calm-v2",
  "format": "audio/l16;rate=24000;channels=1"
}

This is an application record, not a Gemini request schema. It gives generation, caching and incident review the same identity.

Treat streaming audio as a byte protocol

The speech-generation guide documents an important format boundary. A unary request returns a complete WAV file by default, including a RIFF header. A streaming request returns headerless 16-bit signed little-endian linear PCM chunks at 24 kHz mono by default. Concatenating several WAV responses without removing their headers produces a malformed stream.

For browser or app playback, buffer enough audio to avoid underruns but begin only after validating the first chunk and format. Preserve ordering, apply backpressure and stop delivery when the client disconnects. For telephony, explicitly request G.711 mu-law or A-law and the required sample rate rather than transcoding blindly after generation. Store the declared MIME type, sample rate, bit depth and channel count beside every artifact.

A streaming retry needs a new delivery plan. If synthesis fails after the listener heard three seconds, replaying from byte zero may repeat speech. Keep sentence or turn boundaries, acknowledge delivered segments locally and restart at a safe semantic boundary. Do not pretend byte-level continuation is guaranteed when the API only gave a new stream.

Give voice identities a lifecycle

Gemini 3.8 supports four voice paths: 30 featured studio voices, an extended catalog queried through `GET /v1beta/voices`, designed voices created from a natural-language persona, and replicated voices derived from reference audio plus consent. Voice design returns a reusable `voice_...` ID and a preview for auditioning.

Treat the voice as versioned product configuration. Record owner, allowed locales, approved use cases, creation source, reviewer, status and replacement voice. A marketing voice may be approved for published narration but not for transactional calls. Store only a logical alias in content templates; resolve it to the current provider ID at synthesis time. That allows revocation or rotation without rewriting every document.

Provider retention is not your governance policy. Google documents a shared limit of 200 stateful custom voices per project and a one-year retention window that resets with use. Stateless replicated `voicekey_...` values are client-managed and expire after seven days. Build deletion, inventory and expiry jobs explicitly instead of discovering missing voices during a live request.

Bind replication to recorded consent and purpose

Google requires two recordings from the same adult speaker for voice replication: 10–30 seconds of clean reference speech and a recording of the mandatory consent statement in a supported locale. The documentation includes an Arabic consent phrase and recommends 24 kHz mono 16-bit WAV. This verifies a provider requirement; it does not by itself establish every legal right for every later use.

Keep the consent artifact, its hash, locale, capture time, speaker identity verification, permitted products, expiry and revocation status in a restricted registry. Separate access to biometric source audio from access to synthesize with an approved voice. Log use of a voice without logging sensitive audio by default. On revocation, disable the logical alias, delete provider-stored identities where required, invalidate caches within policy and preserve only the audit evidence your legal basis permits.

Never accept a public video clip as proof of permission. Do not let one tenant reference another tenant's `voice_...` ID. A voice identifier is a sensitive capability and must be resolved under authenticated tenant scope.

Design multi-speaker output around documented limits

A single multi-speaker request supports up to two speakers using prebuilt voices. Every turn must identify its configured speaker. If the dialogue uses designed or replicated voices, Google's guide says to synthesize each speaker turn separately. Those outputs then need deliberate assembly.

Keep one manifest containing turn number, speaker alias, transcript hash, voice version, style, generated artifact and verified duration. Request raw PCM or strip individual WAV headers before concatenation. Add silence or overlap only through an explicit editor step. This avoids a common anti-pattern: treating several independent files as though they were one continuous container.

Backchannels and overlap tags can improve a performance, but they make timing and accessibility harder. Keep a clean transcript, speaker labels and captions as separate source artifacts. Do not derive captions from the final mixed audio when the approved text already exists.

Cache deterministic audio, not authorization

For stable material, application-level audio caching can remove repeated synthesis. Build the key from transcript hash, model version, resolved voice version, style version, locale and output format. A text edit, voice rotation or changed pronunciation dictionary must produce a different key. Do not cache only by raw text.

Check voice authorization before serving a cached artifact. Revoking a replicated voice may require stopping future playback even when the audio file already exists. Separate immutable object storage from the mutable delivery policy. Use signed, short-lived download URLs for private audio and set lifecycle rules for abandoned drafts.

Do not cache personalized or sensitive output unless the business purpose, access model and deletion policy justify it. Sometimes regeneration is cheaper than building a risky shared cache.

Evaluate the complete listening experience

A single pleasant sample is not an evaluation. Build a suite covering names, Arabic and English numerals, currency, acronyms, addresses, product codes, long paragraphs, emotional changes, punctuation, dialects and malformed input. Include human review by fluent listeners for the actual locales.

Measure text fidelity, pronunciation error, speaker consistency, unwanted omissions or additions, style adherence and listener preference. Operationally track time to first audio, total synthesis time, stream underruns, retries, bytes delivered before failure, cache hit rate, cost per finished minute and completion rate. For replicated voices, add unauthorized-use attempts and revocation propagation time.

Compare Flash with Flash-Lite on the same scripts and voice configuration. Promote a route only when its benefit survives blind review and production-like network conditions. Provider claims about fidelity or efficiency are starting hypotheses, not results from your users.

When not to use generative TTS

Use prerecorded human audio when wording is fixed, emotion is critical and the content changes rarely. Use conventional deterministic TTS when predictable pronunciation, offline execution or a narrow approved voice set matters more than expressive generation. Use the Live API for true bidirectional audio conversation. Do not use replicated voices when consent, revocation and tenant isolation cannot be enforced.

Avoid generative speech for legal, medical, financial or emergency instructions without a human-reviewed transcript, clear provenance and a safer fallback. Do not generate a voice to impersonate a public figure or a colleague. And do not turn private text into audio merely because the interface supports it; apply the same data-classification policy used for the original content.

Production rollout checklist

Start with prebuilt voices and non-sensitive, short, read-only content. Lock transcript normalization and output formats, then test unary and streaming paths separately. Add Flash-Lite as the default route and compare Flash only where the evaluation shows a material quality gain. Introduce designed voices after ownership and versioning exist. Enable replicated voices last, behind consent evidence, tenant isolation, revocation and audit.

Run canaries by locale and use case. Define rollback thresholds for pronunciation regressions, missing words, voice drift, first-audio latency, stream failures, unexpected cost and consent-policy violations. Keep a prebuilt fallback voice that does not depend on a custom identity.

The architectural rule is simple: the transcript is the approved message, the voice is an authorized identity, and the audio file is a versioned artifact. Gemini 3.8 TTS can synthesize the performance, but the application must own permission, lifecycle, delivery and truth. Connect this pipeline to Gemini 3.8 Live voice-agent architecture, durable AI agent execution and backend observability.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceAI integrations & retrieval systemsRelevant projectAI Action Studio