Tavus Says Its New AI Model Passes a Video Turing Test
Tavus (tavus.io) has introduced Griffin, a real-time AI model that holds face-to-face video conversations. In a live study run by the company, 26 of 54 people (48%) who spent one minute on a video call with a Griffin-powered character came away believing they had talked to a real person. Participants were told their partner was another study participant. When Tavus ran the same test on its previous production system, 1 of 41 people said their partner was human.
Tavus calls Griffin its first "Human Interaction Model." The company was founded by Hassaan Raza and Quinn Favret and went through Y Combinator in Summer 2021. Its backers include Sequoia Capital, which led its seed round, and CRV, which led its $40 million Series B last November. Tavus builds AI characters called PALs that people can see and talk to. Developers can access them through an API, and consumers can talk to them directly.
Most voice and video assistants work in turns. They wait for the user to stop talking, then run the audio through speech recognition, a language model, speech synthesis, and finally an avatar renderer. That pipeline creates the familiar pause before the face moves, and it often mistakes a pause for thought as the end of a turn. Griffin replaces the pipeline with one system that watches and listens continuously. Several times a second, it decides whether to stay quiet, nod, say "mm-hm," cut in, or yield when interrupted.
The model has two parts. The first is a conversational engine that takes in the user's audio and video and outputs control signals covering what to say and how to say it, including tone, facial expression, and gesture. The second turns those signals into speech and video together. Because both parts run continuously, Griffin can change its face or voice in the middle of a sentence when the person on the other end reacts.
"In great conversations, you aren't thinking about the conversation at all."
Hassaan Raza, Co-founder and CEO of TavusGriffin also uses vision for more than reading faces. In demo clips released with the research post, it coaches someone through a Rubik's Cube while watching their hands, and it guides an engineer through soldering a board, speaking up when the next step is due. Each frame is generated in real time from a single reference image, including the body, the chair, shadows, and the background. Tavus says the video comes out at 720p in 320-millisecond chunks, with average audio-to-video latency of 0.43 seconds on Nvidia H100 GPUs.
Outside the company's own study, Tavus points to NVIDIA's VideoFDB benchmark for full-duplex audio-visual conversation, which NVIDIA scored independently in September. On the benchmark's generation track, Griffin scored 3.83 out of 5. The human reference scored 3.92, and the next-best system scored 2.80. Griffin also placed first on the perception track.
The research was led by Raza and Ioannis Patras, Tavus's head of research. Tavus has not released Griffin widely. A smaller version, Griffin-Lite, is available as a research preview to a limited group of testers, and a more capable model will follow. The company says a model that people can mistake for a human needs more safety and disclosure work before a broader release.