AI Models & Platforms

Google Brings Live Avatar Visual Presence to Gemini 3.8 Live

mm
Add Unite.AI to your preferred sources on Google

Google on September 24, 2026 introduced Gemini 3.8 Live with Live Avatar, a feature that adds near real-time generated video to the speech of its native live dialogue models, and made it generally available in Gemini Enterprise with endpoints in the United States and European Union.

The announcement, written by research scientist Shuo-yiin Chang and software engineer CJ Zheng on behalf of the Gemini Audio Team, presents the feature as a visual presence layer for Gemini’s conversational AI. Google says its precise lip-syncing, natural expressions, and fluid turn-taking let enterprises expand virtual offerings such as customer service and interactive walkthroughs.

The release follows Google’s September 15, 2026 launch of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, live dialogue models Google positioned for scalable, cost-efficient conversation and for high-complexity reasoning tasks respectively, which began rolling out through the Gemini API, Google AI Studio, the Gemini app, Google Workspace, and Search Live.

General Availability in Gemini Enterprise

Gemini 3.8 Live with Live Avatar is generally available in Gemini Enterprise as of September 24, 2026, and Google Cloud said the technology was first previewed at Google Cloud Next 2026. The release is offered through US and EU endpoints and includes provisioned throughput, enterprise compliance, and strict data governance, according to the Google Cloud post by Fabien Blanc-paques, group product manager for Gemini Live. Gemini 3.8 Live Extended Thinking remains in private preview.

Enterprise customers can try the model in the Gemini Enterprise console, integrate it through the Gemini Live API documentation, and review pricing on Google Cloud’s pricing page. Google Cloud tells customers to work with its sales representatives for provisioned throughput, customized deployment architectures, and custom-avatar allowlisting.

How Live Avatar Works

Live Avatar processes visual and audio inputs at the same time and answers with expressive audio and video in near real time, according to Google. The feature also supports asynchronous tool calling, so it can start tool calls and retrieve data in the background without pausing the conversation. One Google demonstration shows the avatar checking in a hotel guest while the dialogue continues without interruption.

Google says the avatar adjusts its lip-sync and expressions as conversations move across 97 languages, maintaining video fidelity and avoiding visual drift. The Google Cloud post adds live visual understanding of camera feeds and screen shares alongside audio, interruption recovery that preserves conversation context and backend transactions, automatic language detection and transition without manual toggles, and deployment across web, mobile, and interactive kiosks. A separate Google Cloud demonstration shows an insurance claims intake agent filling in a claim notebook as a customer shows damage on camera, while an agent team built with Google’s Agent Development Kit verifies the policy, applies the intake rules, and assembles the adjuster packet in the background.

Avatars, Watermarking, and Model-Card Limits

Organizations can deploy from a library of preset avatars or generate custom avatars from a high-quality reference image, and Google says the result retains the likeness, brand styling, or character identity of the reference. Access to custom avatar creation requires enterprise allowlisting; Google Cloud describes the allowlisting and verification process as a safeguard for identity and against misuse. Every generated audio and video stream is embedded with an imperceptible SynthID watermark, keeping the AI-generated content detectable.

According to the Gemini 3.8 Audio model card, the models are based on Gemini 3 Pro. Gemini 3.8 Live accepts audio, images, video, and text with a context window of up to 128K tokens, and with Live Avatar it outputs audio, video, and text subject to a 24K-token output limit. The model card states that the Live Avatar variants produce expressive video of natural head movements synchronized to speech, driven by configured images and audio inputs, and that they sustain a few minutes of continuous interaction rather than extended hours.

The model card lists known limitations including possible hallucinations and occasional slowness or timeouts, and it sets the knowledge cutoff at January 2025. Its frontier safety assessment found no meaningful new capabilities or material performance gains relative to Gemini 3.7 Flash, and on that basis Google considers the Gemini 3.8 Audio models unlikely to reach any Tracked or Critical Capability Levels under its Frontier Safety Framework.

Customer Deployments

Google Cloud named several enterprises building with the feature. Cox Automotive built an AI-powered shopping assistant for Autotrader that uses live screen-highlighting and tool calling to guide car shoppers through vehicle search, comparison, and financing in real time.

“Shoppers increasingly expect to describe what they need in their own words rather than work through filters and menus. Autotrader’s new conversational AI Avatar brings that experience to vehicle discovery by matching natural conversation to the right inventory,” Marianne Johnson, executive vice president and chief product officer of Cox Automotive, said in the Google Cloud announcement.

Equal AI chief executive Akhilesh Damaraju said the company’s personal AI handles more than a million live calls daily across nine Indian languages, and said Gemini 3.8 Live improved interruption handling, multilingual conversations, and tool-call reliability. Bob Van Osten, Salesforce’s vice president of product for Agentforce, said Agentforce and Gemini 3.8 Live are coming together in a collaboration between Salesforce AI Research and Google that combines real-time multimodal capabilities with agentic AI. Eric Walsh, a software engineer at Specs, said the model’s Voice Activity Detection updates and latency improvements were steps forward for the AI assistant on the SPECS platform.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.