Technology

Google Adds Live Talking Avatars to Gemini Enterprise

Martin HollowayPublished 25m ago4 min readBased on 8 sources
Reading level
Google Adds Live Talking Avatars to Gemini Enterprise
source:blog.google

Google has made Gemini 3.8 Live with Live Avatar generally available to Gemini Enterprise customers, adding an animated persona that speaks and reacts in real time during live conversations. The company detailed the launch on September 24, 2026, in a post titled 'Introducing Gemini 3.8 Live with Live Avatar' Google. The change is straightforward. Users can now watch the model respond while they talk with it.

Live avatars create real-time video of a talking face matched to synthesized speech from the gemini-3.8-live model Documentation. The connection runs through the Gemini Live API, which handles low-latency voice and video exchange with Gemini. Low latency means short delay between speaking and response. For companies building assistants, the surrounding tools sit in the Gemini Enterprise Agent Platform, a system to build, scale, govern and optimize agents.

The focus is conversational presence, the sense that someone is listening and responding. The avatar lip-syncs and shows different facial expressions during conversations The Verge. Google says the system also produces natural head movements. Turn-taking is fluid, and visual input is processed in near real time. The result feels closer to a video call than to a step-by-step prompt and answer cycle.

Google describes Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as its most advanced live dialogue models so far, built for natural conversation. Live Avatar is the visual layer on top of that dialogue system. It does not replace text or voice output. It adds a synchronized face to them.

Deployment is enterprise-first. Availability is currently limited to Gemini Enterprise customers. Google is offering the service with US and EU endpoints and with provisioned throughput Google Cloud. Endpoints set where requests are processed, which helps with data residency rules about where data is stored. Provisioned throughput means reserved capacity, which helps with planning and keeps inference latency, or response delay, more consistent under load.

Two other enterprise features are language coverage and identity control. Live Avatar supports 97 languages and can switch between them without loss of video quality or visual drift, where the face slips out of sync. Google will offer a library of preset avatars and let organizations create their own. In practice, a support agent, sales assistant, training coach or internal copilot can keep the same visual identity across markets without rebuilding the video setup for each language.

Origin is marked in the output itself. Live Avatar video includes an invisible SynthID watermark, a hidden signal embedded in the file. The mark travels with the generated video and is intended to allow later identification as synthetic media.

The broader context for teams shipping agents is that the technical change is narrow but practical. A steady, lip-synced face reduces ambiguity in voice-first exchanges. Users get timing cues, emphasis and pauses that audio alone does not always carry. For multilingual deployments, holding video quality steady across language switches avoids a common failure where the voice switches cleanly but the face lags, stutters or resets.

In my view, the enterprise framing is deliberate. Consumer avatars invite short-term novelty use. Enterprise avatars invite measurement. Call containment, task completion, handoff rate and user trust can be tested against text-only and voice-only baselines. Google says Live Avatar enables enterprises to expand their virtual offerings, and that wording points to where early use will likely concentrate. It is customer-facing roles with high volume and repeatable structure, where a consistent persona lowers operational cost without removing human escalation.

For architects weighing a rollout, the practical point is operational load. Real-time video matched to synthesized speech is sensitive to jitter, packet loss and tail latency, or uneven delays in data delivery and slow responses at the edge of the range. Provisioned throughput and regional endpoints help, but teams will still need to track conversation quality, fallback behavior and watermark checks. The avatar makes the system easier for users to read. It also makes failures easier to see.

The longer history here is encouraging. Interfaces have moved from command lines to windows to touch to voice. Each step lowered the effort needed to state intent. We have seen this pattern before, when the command line gave way to the graphical desktop and computing reached more people. A responsive face that listens, speaks and holds context continues that direction. The technology will improve human work when it is governed well, measured honestly and used where presence helps people finish something.