Google introduced Gemini 3.8 Live and Live Extended Thinking on September 15, 2026, then added Live Avatar on September 24. The shared direction is conversational continuity: an agent can keep interacting while reasoning or tool work happens in the background.
The model announcement distinguishes a scalable dialogue model from a version for more complex reasoning. The Avatar release adds streaming video presence and states availability in Gemini Enterprise. Those are related capabilities with distinct deployment surfaces.
Background work changes the conversation contract
Google describes asynchronous tools and a model that can acknowledge a request while continuing a conversation. For a voice agent, this can avoid dead air during a lookup or a multi-step task. It also creates a state-management problem: the user may change a request before the original work finishes.
A production application should define when work is cancelled, when results are discarded, and when the user must reconfirm an action. If a caller changes a delivery address during a lookup, the final write must use the confirmed current address. Fluency should not hide an obsolete tool result.
Extended thinking and visual presence serve different needs
Extended Thinking targets more demanding work. Live Avatar pairs dialogue with near-real-time visual output. An expressive face can make a service easier to follow, but it does not make the underlying answer more accurate.
Evaluate whether video actually helps the task. An instructional walkthrough may benefit from visual presence; a telephone support line cannot use it. A team should account for video delivery, accessibility, network conditions, and disclosure that the user is interacting with AI rather than assume every voice interface needs an avatar.
Measure completed tasks, not just speech quality
Google reports benchmarks and multilingual capabilities. Those are useful starting points for evaluation, with the settings and access path stated in the source. They do not establish a success rate for a business’s own tool chain.
Test interruptions, corrections, slow backends, ambiguous names, and tool failures. Record whether the agent accurately explains what has finished and what is still pending. For a reservation or account change, the result should be independently visible in the underlying system rather than inferred from a confident spoken response.
Where the releases fit in a voice architecture
Teams considering native speech models should compare them against their current pipeline using the same task set. Include latency to a useful answer, total session cost, successful tool outcomes, and user recovery from errors. A model that sounds natural but mishandles a correction may be a worse operational choice.
Nerova’s assessment is that continuous dialogue is meaningful when it remains synchronized with real work. The September releases broaden the options for that experience. Permission checks, transaction ownership, and truthful status reporting still belong in the application surrounding the model.