Google's Live Avatar keeps talking while it fetches data

In Google's hotel check-in demo, the onscreen receptionist never goes quiet. While tool calls pull the booking details in the background, the avatar keeps talking to the guest. The feature is Gemini 3.8 Live with Live Avatar, and Testingcatalog reports it is now generally available to enterprise customers in Gemini Enterprise.
At a glance
- Google's Gemini Audio Team built Live Avatar as an extension of Gemini 3.8 Live, pairing the model's native dialogue with low-latency video so an agent can listen, see and speak through an onscreen persona.
- Google says lip-sync and facial expressions stay synchronized across 97 languages, and that the avatar can switch languages mid-conversation without losing video fidelity or showing visible drift.
- Preset avatars are open to Gemini Enterprise customers, but a branded character built from a reference image currently requires enterprise allowlisting, and every output carries a SynthID watermark in audio and video.
If you missed it, Gemini 3.8 Live shipped only last week as Google's live conversational model, the one that handles spoken dialogue as it happens. Live Avatar gives that same model a face. Google Cloud Tech announced general availability in Gemini Enterprise on September 24, 2026, and invited developers to start building.
Live Avatar keeps talking while tool calls run in the background
Tool calling sits on Google's list of key capabilities, and it runs asynchronously. Live Avatar can call tools and retrieve data while it goes on speaking, so an enterprise agent can finish a task without dropping out of the conversation. In the hotel check-in demo, the persona stays on screen and engaged while the information needed for check-in is fetched behind the scenes.
The other demonstrations show several different avatar characters, a conversation running in near real time, and a single exchange in which the language switches midway. Google positions the product for customer service and guided walkthroughs. In each case the agent takes audio and visual input at the same time, so it can hear and see the person in front of it while it talks.
Google says the lip-sync covers 97 languages
The persona is meant to make a conversation feel closer to a face-to-face exchange. Google lists precise lip-syncing, natural facial expressions and fluid turn-taking, the back-and-forth in which each side knows when to speak and when to listen. The Google Cloud Tech announcement names video avatars, fluid dialogue and tool calling among the headline capabilities.
Google says its native multilingual speech-to-speech synchronization covers 97 languages. When the spoken language changes, the lip movements and expressions change with it. According to Google, the system is designed to move between languages without any loss in video fidelity and without the face visibly drifting out of sync with the voice.
Branded avatars from a reference image need enterprise allowlisting
Companies can pick from a library of preset avatars or build a branded character from a high-quality reference image. Google says the result can keep the source likeness, visual styling or character identity while becoming fully animated and responsive. Custom creation is more restricted than the core product and currently requires enterprise allowlisting.
The release ships with safeguards that Google describes as intended to respect identity. SynthID watermarking is embedded in both the audio and the video output, and Google says the imperceptible marker is meant to keep AI-generated material detectable and to reduce misinformation or misattribution. For deployment guidance, Google points enterprise developers to its model card and API documentation.
The Gemini Audio Team coupled dialogue and video inside Gemini 3.8 Live
According to Google, Live Avatar natively couples the model's dialogue capabilities with low-latency video. Put simply, the system that decides what to say is also the one driving how the face says it, instead of passing finished audio along to a separate animation step. That is why lip movements can follow the speech when the language changes mid-sentence.
Tool calling works in the same spirit. A request to a booking system or database runs asynchronously, so the conversation does not have to stop and wait for the answer. Picture two phone agents: one says "please hold" and puts on music while checking your reservation, and the other keeps chatting while typing into the system. Live Avatar is built to behave like the second one.
The announcement leaves a lot out. Google gives no latency figures beyond "near real-time" and "low-latency", no pricing and no list of the 97 languages, and the fidelity and drift claims rest on Google's own demos. In our view, keeping branded avatars behind allowlisting while the presets are generally available is the right order of release for a tool that can animate a likeness from one reference image.
When branded avatars open up
Google has given no date or criteria for moving custom avatar creation beyond enterprise allowlisting. For now the practical checkpoint is the documentation. The model card and API docs are where enterprise developers will learn what deploying Live Avatar actually involves. The first customer-service rollouts will show whether the turn-taking and language switching hold up outside Google's demos.
Related stories
- Gemini 3.8 Live Extended Thinking tops GPT Live 1
- Gemini 3.8 Flash TTS takes stage directions line by line
- Gemini Notebook reports generate slides on demand
- Custom MCP servers reach Gemini Business in preview
- Lyria 3.5 lands in the Gemini app, AI Studio and APIs
- Gemini Live adds voice commands for Workspace tasks
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
