Our visual story this week comes from a one-minute video call. Tavus has launched Griffin, a real-time video model that sees, listens and speaks at the same time, and that reacts with its face while the other person is still talking. In the company's own study, 26 of 54 participants, or 48%, believed they had been speaking to a real human. Tavus says earlier systems managed no more than 2%.
The important change is timing, not intelligence. Older avatars chained together speech recognition, a language model, text-to-speech and lip-sync, and the gaps between each step gave them away. Griffin runs one continuous loop, with a reported latency of 0.43 seconds, so the pauses, nods and interruptions arrive when a human would expect them. People judge a caller on these small cues far more than on the content of what is said.
Tavus is restricting access to trusted testers while it builds disclosure measures, but we can expect AI avatars to get a good deal better soon.
