The video Turing test asks one question: on a live video call, can a person tell whether the other side is a human or an AI? It takes Alan Turing's 1950 imitation game, which was text only, and moves it to a face, a voice, and real-time timing. It is back in the news because on October 1, 2026 Tavus said its new model, Griffin, "passed" it. This explainer covers how the original test works, how the text version fell, what changes with video, and how to read Tavus's number without over- or under-believing it. For the launch details and benchmarks, see our Tavus Griffin news post.
TL;DR: the video Turing test in one table
| Question | Short answer |
|---|---|
| What is it? | A live video-call version of Turing's imitation game |
| Who sets the pass mark? | Nobody. It depends on the experiment's design |
| Original test medium | Text only (1950) |
| Text test result | GPT-4.5 with a persona was judged human 73% of the time in a five-minute three-party test (UC San Diego) |
| Tavus claim | 26 of 54 testers (48%) said a one-minute video partner was human; prior stack scored 1 of 41 (2.4%) |
| Biggest caveat | Short, company-run, single-partner protocol; no independent replication |
| Why it matters | It measures how easily a face and voice can be trusted by default |
| What to do | Disclose AI on calls and use out-of-band verification for money |
How does the original Turing test work?
In his 1950 paper, Turing replaced the vague question "can machines think?" with a game. A human interrogator talks by text to two hidden parties, one human and one machine, and must decide which is which. If the machine is hard to distinguish, it has shown a behavior we would call intelligent.
Turing did not set a strict pass mark, but he made a prediction: that in about fifty years an average interrogator would have no more than a 70% chance of making the right identification after five minutes of questioning. That implies roughly 30% of interrogators fooled, which is where the often-quoted "30% rule" comes from.
Two design features matter for everything that follows:
- Three parties. The classic setup shows the interrogator a human and a machine side by side, so a judge can compare.
- Free questioning. The interrogator chooses what to ask, which lets a skilled judge probe weaknesses.
For the longer history of the idea and why it is contested as a measure of intelligence, see our history of AI and AI consciousness guide.
Has any AI passed the text version?
By one prominent measure, yes. Cameron Jones and Benjamin Bergen at UC San Diego ran two pre-registered, randomized three-party Turing tests with five-minute conversations. When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time, significantly more often than interrogators picked the real human. LLaMa-3.1-405B with the same prompt reached 56%, while ELIZA and GPT-4o without the persona scored 23% and 21%.
The lesson: passing depends on prompt, persona, and setup. The same model without a persona failed. A Turing result is a property of the system plus the instructions plus the experiment, not of the model alone. Keep that in mind for the video version.
What does "video" change?
Moving to video adds cues that text did not have, and each one is a place an AI can slip or succeed.
| Cue | Why it matters on video |
|---|---|
| Face and lip sync | Mismatch between audio and mouth is a classic tell |
| Voice prosody | Tone, breathing, and filler words carry a lot of trust |
| Turn-taking | Humans interrupt, overlap, and react in under a second |
| Gaze and listening | A partner that looks at you and reacts while you speak feels present |
| Latency | Delays of even a few hundred milliseconds feel off |
| Perception | Noticing someone entering the frame or a gesture requires real-time vision |
This is why Tavus frames Griffin as a "Human Interaction Model": a full-duplex video-to-video system that listens, watches, and responds at the same time rather than chaining speech recognition, a language model, text-to-speech, and an avatar. Older cascade systems add latency at every handoff, which is exactly what people detect. The comparison between cascades and unified models is the same one we covered for voice in GPT-Live-1 and for avatars in HeyGen LiveAvatar.
Video also changes human psychology. People extend default trust to a face and a natural voice, so a video partner has to cross a lower bar to pass than a text bot facing a suspicious interrogator.
What exactly did Tavus test?
Per Tavus's October 1, 2026 write-up, as summarized in our launch post:
- Participants came through an independent research platform.
- They were told they would be matched with another participant for a one-minute call about what they were looking forward to this year.
- The partner was actually a Griffin-Lite AI that generated the face, voice, and replies live. A comparison arm used Tavus's earlier stack.
- After the call, participants rated naturalness, trust, flow, and whether they felt listened to.
- Only at the end were they asked whether it had crossed their mind that the partner might be an AI. Then everyone was told.
| Measure | Griffin-Lite (n=54) | Earlier Tavus stack (n=41) |
|---|---|---|
| Said "real person" | 26 of 54 (48%) | 1 of 41 (2.4%) |
| Doubters' first suspicion | Usually under 20 seconds | n/a |
| Conversation flow (7-point) | 4.9, the lowest rating | n/a |
The jump from 2.4% to 48% is large. The question is what it means.
How should you read the 48%?
Four design choices make this a different experiment from a classic Turing test.
- Single partner, no comparison. Participants judged one partner, not a human next to a machine. With no human to compare against, "real" is the default assumption on a video call.
- Deceptive framing. They were told their partner was another participant. In Jones and Bergen's design, people knew one of two partners was a machine and were actively hunting for it. Suspicion changes behavior a lot.
- One minute, one topic. A friendly small-talk prompt is far easier than five minutes of adversarial questioning.
- Detection was an afterthought. Participants were asked whether the AI possibility had crossed their mind only at the end. Over half said it never did.
So 48% is not "half of people can't tell on any call." It is "about half did not suspect an AI during a one-minute friendly chat when nobody told them to look." That is still a meaningful finding, because it describes the real-world situation of an unexpected video call. But it is a measure of default trust and a short protocol, not of an interrogator's best effort.
It also helps to compare against Turing's own bar. If you use his rough 30% figure, 48% is well above it. If you use 50% as chance in a three-party test, 48% looks like chance-level indistinguishability, but that framing only holds when a human is also on screen, which was not the case here. The two framings are why you should be wary of anyone declaring a clean "pass."
What are the other measures besides a Turing test?
Tavus also cited NVIDIA's Video Full-Duplex Benchmark, where Griffin-Lite scored 3.83 on generation against 3.92 for real humans, and 3.73 on perception. A benchmark like that scores specific behaviors (turn-taking, expressiveness, noticing visual events), which makes it more diagnostic than a single yes-or-no judgment. A Turing-style number tells you whether people were fooled; a benchmark tells you why. For serious evaluation, ask vendors for both, plus the sample size.
Why does this matter outside the lab?
Because the capability being measured is also the capability that enables fraud and manipulation. A model that holds a believable live conversation with a generated face has obvious uses for support, coaching, sales, and tutoring, and equally obvious uses for impersonation. Tavus itself says Griffin-Lite is limited to trusted testers because the same property "looks like a product" and "looks like deception," and it is working on disclosure features.
Real incidents show why. The $25.6 million deepfake video-call fraud worked because employees trusted faces they saw on a call. And as we asked in are we all meat proxies now, the more natural the interface, the more people stop checking who or what is on the other end.
What should you do about it?
- Disclose AI. If you deploy an avatar, say so in the first sentence of the call and in the UI.
- Verify out of band. For payments, credentials, or approvals, confirm through a separate channel you initiated.
- Evaluate with a protocol. If you test a vendor, run a longer call, tell testers to look for an AI, include a human comparison, and record sample sizes.
- Ask for scores, not slogans. Prefer benchmark tracks and counts over percentages-ahead marketing.
- Do not hang a roadmap on a testers-only model. Griffin-Lite is a research preview, not a customer product.
Questions people are asking
Is the video Turing test the same as detecting deepfakes?
Not quite. Deepfake detection analyzes artifacts in media. The video Turing test is a human-judgment experiment about live interaction. A system could be flagged by software and still fool most people, or the reverse.
Does passing mean the AI is intelligent?
No. Turing's test is behavioral. It shows that an interaction was convincing, not that the system understands anything. Many researchers argue the test mostly measures how easily humans anthropomorphize.
Will there be a standard video Turing test?
Unlikely soon. Different groups use different durations, framing, and scoring. Expect independent replications with longer calls and explicit-suspicion conditions before any consensus emerges.
Honest limitations of this post
- Tavus's numbers are company-reported and come from the launch write-up as summarized in our earlier post. We have not seen independent replication.
- We did not run our own test. Statements about what the protocol implies are analysis, not new data.
- Text-test figures come from the UC San Diego study as reported; check the paper for exact methods.
Related reading
- Tavus Griffin: a video Turing test, not a ship
- Are we all meat proxies now?
- Deepfake fraud: the $25.6 million video call
- HeyGen LiveAvatar and GPT-Live-1 demos
- GPT-Live-1 in the API
- History of artificial intelligence, 1950-2026
- AI consciousness and sentience guide
- Is AI taking us to a point of no return?
Official: Griffin research post (Tavus) · Large Language Models Pass the Turing Test (arXiv)
Accurate as of October 2, 2026. Tavus figures are company-reported.
