DeepMind has introduced Gemini 3.8 Live with Live Avatar, an update to last week's Gemini 3.8 Live that adds a streaming visual presence to the conversational AI. According to the Gemini Audio Team, the system pairs near-real-time video generation with speech, so the avatar listens, sees, and speaks with lip-syncing, expressions, and turn-taking. It is available in Gemini Enterprise starting today.
DeepMind frames conversation as inherently multimodal, and Live Avatar is designed for enterprise use cases such as customer service and interactive walkthroughs. The underlying model processes visual and audio inputs simultaneously, and can execute asynchronous tool calls in the background while continuing the dialogue. The company says this lets the avatar handle complex tasks, such as checking in a hotel guest, without interrupting the flow of conversation.
Live Avatar supports native multilingual speech-to-speech synchronization across 97 languages, with lip-sync and expressions adapting mid-conversation. Organizations can choose from preset avatars or generate custom ones from a reference image, though custom creation is currently available only through enterprise allowlisting. All output is watermarked with SynthID to keep AI-generated content detectable, and DeepMind points to its model card for safety details.