How Believable is Google's New 'Live Avatar' Capability?

by

Google has synthesized "expressive face-to-face experiences" for its speech agent Gemini 3.8 Live. They're now offering a Live Avatar "with precise lip-syncing, natural expressions, and fluid turn-taking" for Google Enterprise accounts wanting "engaging customer service" or for offering interactive walkthroughs. (Check out the not-creepy-at-all video in Google's announcement.)

"Though Google will offer a library of preset avatars for customers to choose from, it will also allow organizations to create their own," notes The Verge. (See some examples from the YouTube channel "AI with Surya".)

But even without the visualization of the avatar, "I was never able to shake the feeling that these conversations with computers never feel like a real conversation," argues the blog Android Police. Conversing with just the Ai-generated audio, "At best, they feel like talking to a phone representative or someone from tech support. We turn to them when we have a problem, and they help us through it..." GPT-Live, and Gemini Live right behind it, skip that whole relay race. Instead of translating your voice to text and back to voice, the model works with raw audio the entire way through (what it hears and what it says) inside the same system, with nothing translated in between. That sounds like a small plumbing detail, but it's the whole story. Cutting out the text step lets these models respond in a fraction of a second instead of the pause we've learned to expect, and it lets them hear things text can never carry: tone, hesitation, whether you're annoyed or joking...

I was hoping to be surprised by how natural the conversation felt. Instead, I came out with a deeper appreciation for every human I've ever talked to. Even the boring ones... I spoke, it spoke back. I spoke faster, it answered faster. Then I switched to a different language, and it switched along with me. Even switching between languages several times during the same sentence didn't stump it. The most impressive moment happened when I asked it what "T-O-P-G-3-3-K" spelled out, and it immediately came back with, "You are spelling the word Top Geek, but using a 3 to represent a reversed E...."

Although it felt fast and responsive, at no point did it feel like talking to another human being... What Gemini couldn't replicate, because it was never built to replicate it, is human connection.... I'm sure I'm not telling you something you don't already know, but somehow talking naturally to an LLM amplifies the feeling that there is no one on the other side of the line. It might flow like a phone call, but it doesn't feel like one.

MrBrklyn (Slashdot reader #4,775) says he discussed "why mainstream media avoids reporting on screen dependency" with Gemini, and eventually convinced Gemini to respond that it's just "another tool built by the same tech giants to make sure you rely on their system to tell you what to think, how to talk, and what is real."