A voice agent is not a chat endpoint with a microphone attached. It is a long-lived distributed session in which audio arrives continuously, the model may speak while tools are still running, the user can interrupt, and a network reconnection must not repeat a real-world action. Prompt quality matters, but the production boundary is the session architecture.
Google's Gemini API changelog records the general availability of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026. This article is therefore a current-documentation explainer, not a claim that the models launched today. The standard Live model is the default option for low-latency dialogue and supports interleaved reasoning, native audio and asynchronous function calling. Extended Thinking adds background reasoning for more complex work while continuing the audio session.
Both model pages list 131,072 input tokens and 65,536 output tokens, with text, image, audio and video inputs and text or audio outputs. They also list important absences: no context caching, structured outputs, file search, code execution or URL context. A production design must accommodate those boundaries instead of hiding them behind a clever system prompt.
Choose the model by turn budget
Use Gemini 3.8 Live when the conversation itself is the critical path: support triage, a guided form, a store assistant or a dispatch workflow where the first audible response and natural interruption matter. Use Extended Thinking when a turn needs genuinely deeper background reasoning, such as comparing several policy constraints before proposing the next action. Extended Thinking supports asynchronous function calls only; it is not a universal quality switch for every greeting or lookup.
Route by the task, not by the prestige of the model name. A practical agent can keep ordinary turns on the standard model and hand a bounded complex decision to another service or a separate reasoning step. Measure time to first audio, completed-task accuracy and cost per resolved session. An extra second of reasoning may be useful for a plan, but harmful when the user only asked for an order status.
The published capability table also means that retrieval is an application responsibility. If the agent needs a private knowledge base, expose a narrow retrieval function through your tool layer, attach citations to its result and validate the returned shape. Do not describe file search or structured output as native features when the model page says they are unavailable.
Separate the media edge from business actions
The Live API quickstart uses a persistent WebSocket and an asynchronous interface. Keep the realtime media path small: microphone capture, resampling, chunking, playback, interruption and connection state. Google's best-practices guide recommends 16 kHz input audio and chunks of roughly 20–100 milliseconds. Large buffers raise conversational latency; tiny, irregular writes add scheduling and network overhead.
Two client patterns are reasonable. In a server-mediated pattern, the device streams to your backend, which owns the Gemini connection. This centralizes policy and telemetry but adds another network hop and media-scaling burden. In a client-to-server pattern, the device connects directly to Gemini. Google's ephemeral-token documentation recommends short-lived, constrained tokens for this approach so a long-lived API key is not shipped in a browser or mobile application. The feature is documented as Preview, so lock tokens to the model and session configuration, authenticate issuance on your backend and keep a server-mediated fallback if your risk profile requires it.
Whichever media path you choose, keep privileged business tools behind your backend. A browser may carry audio directly, but it should not hold credentials that can refund orders, change bookings or read another tenant's account.
Treat asynchronous tools as durable commands
For Gemini 3.8 Live, non-blocking function calling is the default. The capabilities guide documents scheduling modes for the standard model, while Extended Thinking accepts asynchronous calls only. The client must still execute each function and send the tool response; the Live API does not turn a declared function into a reliable workflow engine.
Put a tool broker between the model session and business services. Validate the function name and arguments against an allowlist, bind the authenticated user and tenant on the server, and require confirmation for expensive or irreversible actions. Give every call a durable identity such as session ID plus model call ID plus tool name. Store its state before dispatch.
received -> authorized -> executing -> succeeded -> response_sent
\-> unknown -> reconcile -> succeeded or reviewUse an idempotency key when the downstream service supports one. When it does not, add an application-side command table and a reconciliation read. A timeout after submitting a booking is an unknown result, not permission to create a second booking. Replaying a harmless search is different from replaying a payment or cancellation.
Let the conversation acknowledge a long action without inventing its result. The model can say it is checking availability while the tool runs, but the final confirmation must be grounded in a recorded tool response. If the user interrupts or changes intent, cancel only work that is actually cancellable; otherwise preserve the result and explain its state.
Make interruption and reconnection normal states
Voice users talk over the agent. The Live API reports interruptions, and client content can explicitly terminate a generation. On interruption, stop local playback immediately, discard queued audio that the user should no longer hear, and record which model turn was truncated. Do not assume that stopping audio cancels a tool already dispatched. The UI and session state should distinguish “speech interrupted” from “business action cancelled.”
Connections also end. Google's session-management guide describes resumption tokens delivered in `SessionResumptionUpdate` messages. Persist the latest handle and reconnect with it; the best-practices page says handles remain valid for two hours after a session terminates. Handle `GoAway` by using its remaining-time signal to reconnect gracefully instead of waiting for a hard failure.
Resumption preserves model context, but your application still needs durable state: authenticated principal, tenant, conversation purpose, tool-call ledger, pending approval, playback position and the last client event acknowledged locally. Never reconstruct authority solely from the model transcript after reconnecting.
Bound context and listening cost
Audio consumes context quickly. The best-practices guide estimates about 25 tokens per second and says that, without context compression, audio-only sessions are limited to about 15 minutes and audio-video sessions to about two minutes. Configure context-window compression with a trigger and sliding window appropriate to the task. Summarize business facts into your own state before old conversational tokens are evicted.
Billing is token-based. Proactive audio is permanently enabled for both Gemini 3.8 Live models, so input tokens accrue while the API is listening. Input or output transcription adds text-token charges on top of audio. Do not leave an idle connection open indefinitely. End abandoned sessions, pause capture when the product state permits it, set maximum duration and expose the cost impact of optional transcription.
Track audio-in tokens, audio-out tokens, transcription tokens, retained context, silence ratio and cost per completed task. A cheap per-token price can still produce an expensive session when a microphone streams for an hour.
Test the failures that produce duplicate actions
Run a staged test matrix before production: packet jitter, microphone sample-rate mismatch, user interruption during playback, WebSocket reset before and after a tool call, delayed tool results, duplicate function calls, expired ephemeral token, resumption with an old handle, context compression during an active task and a client reconnecting from two devices.
Observe time to first audio, turn latency at p50 and p95, interruption-to-silence time, tool wait, duplicate commands suppressed, unknown commands awaiting reconciliation, reconnect success, context resets, token usage and cost per session. Trace one correlation ID from the user turn through model events, tool command and business result without logging raw sensitive audio by default.
Roll out first to read-only tasks. Then add reversible writes with explicit confirmation, followed by higher-risk actions only when idempotency, audit and recovery have been proven. A production voice agent succeeds when conversation remains natural while state changes remain boringly deterministic. Design the WebSocket, tool ledger, authorization boundary and reconnection path together; the prompt is only one component of that system.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
- Google AI — Gemini API changelog, verified 28 September 2026
- Google AI — Gemini 3.8 Live model, verified 28 September 2026
- Google AI — Gemini 3.8 Live Extended Thinking, verified 28 September 2026
- Google AI — Live API SDK quickstart, verified 28 September 2026
- Google AI — Live API best practices, verified 28 September 2026
- Google AI — Live API session management, verified 28 September 2026
- Google AI — Live API capabilities, verified 28 September 2026
- Google AI — Ephemeral tokens, verified 28 September 2026
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




