On September 23, Google announced Gemini 3.8 Live with Live Avatar, a feature that pairs near real-time video generation with its voice-to-voice conversation models. The feature is available in Gemini Enterprise starting today, with US and EU endpoints, provisioned throughput, and enterprise data governance controls.
The pitch is simple: an AI agent that listens, sees, and speaks with a visual persona. Precise lip-syncing, natural expressions, fluid turn-taking. The kind of thing that sounds like a demo until you look at what it actually solves.
The Feature Is the Interruption Fix
Live Avatar's core capability is not the face. It's what Google calls asynchronous tool execution.
When a user asks for something that requires a data lookup — a hotel check-in, an insurance claim intake — the system triggers the tool call in the background and keeps the conversation going. It doesn't go silent while it waits. The avatar keeps talking, keeps reacting, keeps the presence alive.
That matters for a specific reason. Traditional voice agents pause when they look something up. The silence is awkward, the user wonders if the call dropped, and the experience feels like talking to a system that has to stop thinking to answer. Google's design keeps the reasoning and the conversation running in parallel.
The face is what fills the gap. A talking head with synchronized lips and expressions makes the wait invisible. The enterprise value proposition is not “more human avatars.” It's “fewer dead air moments.”

97 Languages, No Visual Drift
Google says Live Avatar supports 97 languages with speech-to-speech synchronization, switching mid-conversation without “visual drift” — the avatar's lip movements and expressions stay calibrated to whatever language it's speaking, rather than defaulting to one phoneme set.
That detail matters for compliance as much as experience. The EU AI Act, whose transparency obligations took effect August 2, 2026, classifies certain deepfake-adjacent technologies as high-risk and requires disclosure. An avatar that visibly breaks when it switches languages is easier to identify as synthetic. An avatar that doesn't is harder — and therefore more regulated.
Google's answer is SynthID, an imperceptible watermark woven into all generated audio and video. The watermark survives compression and format conversion, and Google says it has been applied to more than 10 billion pieces of content. It's a compliance feature. It's also the only thing standing between “interactive avatar” and “undisclosed synthetic human” in jurisdictions that care about the difference.
The Custom Face Has a Waitlist
Enterprises can choose from preset avatars. To create a custom one — a brand persona, a consistent character — they need to go through Google's allowlisting and verification process.
That gate is deliberate. Google could have made custom avatar generation available to every Gemini Enterprise customer on day one. Instead, it added a verification step that slows adoption.
The logic is identity protection. A world where any enterprise can generate a photorealistic avatar from a reference image is a world where the same capability is available to anyone who wants to impersonate someone. Google chose to slow its own distribution to keep that door narrower.
The cost is adoption speed. The benefit is that when a custom avatar appears in front of a customer, it's attached to a verified organization.
What the Launch Actually Changes
Google is not the first to put a face on an agent. Meta's Muse Realtime Avatar renders a talking character from a reference image, at 448×768 and 25fps. But Meta's version has no API, no price, and no availability date. Google's has model IDs, per-minute pricing, and a GA date.
That asymmetry matters for builders. One is a product you can integrate today. The other is a demo from a keynote.
The enterprise AI conversation is moving toward presence. The question is not whether agents will have faces. It's whether the face is a feature you can trust with a customer, or a liability you have to disclose. Google's answer is: both, and the watermark is the difference.
P.S. Google has not specified what enterprises must tell customers about talking to an avatar. The SynthID watermark is invisible to humans and requires a tool to detect. That leaves a disclosure gap between what the technology can prove and what a user will actually know.
Frequently Asked Questions
Q: What is Gemini 3.8 Live Avatar?
A: A feature in Gemini Enterprise that pairs near real-time video generation with voice-to-voice conversation. It renders a talking digital persona with synchronized lips and expressions during live conversations.
Q: What problem does it actually solve?
A: Interruption. The system runs tool calls in the background while the conversation continues, so the avatar doesn't go silent during data lookups like hotel check-in or insurance intake.
Q: How many languages does it support?
A: 97 languages with speech-to-speech synchronization. Google says switching mid-conversation doesn't cause visual drift — the avatar's lip movements stay calibrated to the language being spoken.
Q: What is SynthID?
A: An imperceptible watermark woven into all generated audio and video. It survives compression and format conversion, and Google says it has been applied to over 10 billion pieces of content.
Q: Can enterprises create custom avatars?
A: Yes, but only through Google's allowlisting and verification process. Preset avatars are available to all Gemini Enterprise customers; custom ones require identity verification.
Q: How does this compare to Meta's avatar?
A: Meta's Muse Realtime Avatar also renders a talking character, but has no API, no price, and no availability date. Google's has model IDs, per-minute pricing, and a general availability date.
