AIToday
Large Language ModelsVideo GenerationAI Safety & AlignmentGIGAZINE AIPublished: Oct 6, 2026, 22:00 JST

Tavus's Griffin-Lite fools 48% in video calls

Tavus's Griffin-Lite fools 48% in video calls

3 Key Points

  1. What happened

    Tavus ran a study where 54 participants had a one-minute video call believing they were talking to another person; 26, about 48%, judged the AI-driven Griffin-Lite to be human, with 79% average confidence.

  2. Why it matters

    If roughly half of participants in a controlled study take an AI for a real person, that is a sign such systems could be misused to deceive people, which is why Tavus is restricting access.

  3. What to watch

    Whether Tavus's safety measures work will determine when Griffin becomes broadly available, and until then it stays a limited research preview for trusted testers.

WHO IT HITSTrust and safety teams evaluating synthetic media, and video-call platforms, may need to assess whether AI participants can pass as human in their products.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Tavus built Griffin as a so-called Human Interaction Model. Unlike typical voice AI that waits for a user to finish before recognizing, generating a reply, and synthesizing speech, Griffin continuously reads video and audio and can nod, murmur, or interrupt in under a second. It generates not just a face but arms, fingers, body motion, chair, shadow, and background from one reference image, at 720p in 320-millisecond chunks, with an average 0.43-second delay from received audio to on-screen video on NVIDIA H100 hardware. Tavus says that is about half the delay of the next-fastest streaming video generation method it compared.

On NVIDIA's VideoFDB benchmark, Griffin-Lite scored 3.83 out of 5 on Generation, close to the 3.92 for human reference data and above Gemini 2.5 + Anam at 2.80. On Perception it scored 3.73 versus 4.20 for humans and 3.44 for MiniCPM-o 4.5; Tavus says NVIDIA ran that evaluation independently in September 2026. In the conversation study, participants rated Griffin-Lite 5.4 for naturalness, 5.6 for trustworthiness, 5.8 for wanting to talk again, and 5.5 for feeling listened to, while natural conversational flow scored lowest at 4.9.

The gap between the 4.9 flow score and the 48% human judgment is the tension Tavus is sitting on. Whether Griffin moves beyond a limited research preview is likely to hinge on how well the company can build visible signs that a caller is an AI without degrading the very naturalness that produced these results.

FAQ
How many participants thought Griffin-Lite was human?
26 out of 54, about 48%, judged it to be a real person, with 79% average confidence.
Why is Griffin-Lite not available to the general public yet?
Tavus says natural human-like conversation could be used to make people think an AI is human, so it is limiting access to trusted testers as a research preview while it works on safety measures.
How does Griffin-Lite compare to Tavus's previous system?
In the same study setup, only 2.4% of 41 participants judged Tavus's previous system, Phoenix-4.5 + Sparrow-2 + Raven-1, to be human.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleWikimedia says OpenAI agents tried to hack its Etherpad tool