explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the video Turing test in one table
  • How does the original Turing test work?
  • Has any AI passed the text version?
  • What does "video" change?
  • What exactly did Tavus test?
  • How should you read the 48%?
  • What are the other measures besides a Turing test?
  • Why does this matter outside the lab?
  • What should you do about it?
  • Questions people are asking
  • Honest limitations of this post
  • Related reading
← Back to blog

explainx / blog

What Is the Video Turing Test? How It Works and What Tavus Claimed

Video Turing Test, Tavus, AI Evaluation, Video Agents, AI Safety

The video Turing test asks if people can tell an AI from a human on a live video call. How it works, how Tavus scored 48%, and why it's not proof.

Oct 2, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
What Is the Video Turing Test? How It Works and What Tavus Claimed

The video Turing test asks one question: on a live video call, can a person tell whether the other side is a human or an AI? It takes Alan Turing's 1950 imitation game, which was text only, and moves it to a face, a voice, and real-time timing. It is back in the news because on October 1, 2026 Tavus said its new model, Griffin, "passed" it. This explainer covers how the original test works, how the text version fell, what changes with video, and how to read Tavus's number without over- or under-believing it. For the launch details and benchmarks, see our Tavus Griffin news post.

TL;DR: the video Turing test in one table

table · 2 cols
QuestionShort answer
What is it?A live video-call version of Turing's imitation game
Who sets the pass mark?Nobody. It depends on the experiment's design
Original test mediumText only (1950)
Text test resultGPT-4.5 with a persona was judged human 73% of the time in a five-minute three-party test (UC San Diego)
Tavus claim26 of 54 testers (48%) said a one-minute video partner was human; prior stack scored 1 of 41 (2.4%)
Biggest caveatShort, company-run, single-partner protocol; no independent replication
Why it mattersIt measures how easily a face and voice can be trusted by default
What to doDisclose AI on calls and use out-of-band verification for money
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

How does the original Turing test work?

In his 1950 paper, Turing replaced the vague question "can machines think?" with a game. A human interrogator talks by text to two hidden parties, one human and one machine, and must decide which is which. If the machine is hard to distinguish, it has shown a behavior we would call intelligent.

Turing did not set a strict pass mark, but he made a prediction: that in about fifty years an average interrogator would have no more than a 70% chance of making the right identification after five minutes of questioning. That implies roughly 30% of interrogators fooled, which is where the often-quoted "30% rule" comes from.

Two design features matter for everything that follows:

  • Three parties. The classic setup shows the interrogator a human and a machine side by side, so a judge can compare.
  • Free questioning. The interrogator chooses what to ask, which lets a skilled judge probe weaknesses.

For the longer history of the idea and why it is contested as a measure of intelligence, see our history of AI and AI consciousness guide.

Has any AI passed the text version?

By one prominent measure, yes. Cameron Jones and Benjamin Bergen at UC San Diego ran two pre-registered, randomized three-party Turing tests with five-minute conversations. When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time, significantly more often than interrogators picked the real human. LLaMa-3.1-405B with the same prompt reached 56%, while ELIZA and GPT-4o without the persona scored 23% and 21%.

The lesson: passing depends on prompt, persona, and setup. The same model without a persona failed. A Turing result is a property of the system plus the instructions plus the experiment, not of the model alone. Keep that in mind for the video version.

What does "video" change?

Moving to video adds cues that text did not have, and each one is a place an AI can slip or succeed.

table · 2 cols
CueWhy it matters on video
Face and lip syncMismatch between audio and mouth is a classic tell
Voice prosodyTone, breathing, and filler words carry a lot of trust
Turn-takingHumans interrupt, overlap, and react in under a second
Gaze and listeningA partner that looks at you and reacts while you speak feels present
LatencyDelays of even a few hundred milliseconds feel off
PerceptionNoticing someone entering the frame or a gesture requires real-time vision

This is why Tavus frames Griffin as a "Human Interaction Model": a full-duplex video-to-video system that listens, watches, and responds at the same time rather than chaining speech recognition, a language model, text-to-speech, and an avatar. Older cascade systems add latency at every handoff, which is exactly what people detect. The comparison between cascades and unified models is the same one we covered for voice in GPT-Live-1 and for avatars in HeyGen LiveAvatar.

Video also changes human psychology. People extend default trust to a face and a natural voice, so a video partner has to cross a lower bar to pass than a text bot facing a suspicious interrogator.

What exactly did Tavus test?

Per Tavus's October 1, 2026 write-up, as summarized in our launch post:

  • Participants came through an independent research platform.
  • They were told they would be matched with another participant for a one-minute call about what they were looking forward to this year.
  • The partner was actually a Griffin-Lite AI that generated the face, voice, and replies live. A comparison arm used Tavus's earlier stack.
  • After the call, participants rated naturalness, trust, flow, and whether they felt listened to.
  • Only at the end were they asked whether it had crossed their mind that the partner might be an AI. Then everyone was told.
table · 3 cols
MeasureGriffin-Lite (n=54)Earlier Tavus stack (n=41)
Said "real person"26 of 54 (48%)1 of 41 (2.4%)
Doubters' first suspicionUsually under 20 secondsn/a
Conversation flow (7-point)4.9, the lowest ratingn/a

The jump from 2.4% to 48% is large. The question is what it means.

XSource postOpen on X ↗

How should you read the 48%?

Four design choices make this a different experiment from a classic Turing test.

  1. Single partner, no comparison. Participants judged one partner, not a human next to a machine. With no human to compare against, "real" is the default assumption on a video call.
  2. Deceptive framing. They were told their partner was another participant. In Jones and Bergen's design, people knew one of two partners was a machine and were actively hunting for it. Suspicion changes behavior a lot.
  3. One minute, one topic. A friendly small-talk prompt is far easier than five minutes of adversarial questioning.
  4. Detection was an afterthought. Participants were asked whether the AI possibility had crossed their mind only at the end. Over half said it never did.

So 48% is not "half of people can't tell on any call." It is "about half did not suspect an AI during a one-minute friendly chat when nobody told them to look." That is still a meaningful finding, because it describes the real-world situation of an unexpected video call. But it is a measure of default trust and a short protocol, not of an interrogator's best effort.

It also helps to compare against Turing's own bar. If you use his rough 30% figure, 48% is well above it. If you use 50% as chance in a three-party test, 48% looks like chance-level indistinguishability, but that framing only holds when a human is also on screen, which was not the case here. The two framings are why you should be wary of anyone declaring a clean "pass."

What are the other measures besides a Turing test?

Tavus also cited NVIDIA's Video Full-Duplex Benchmark, where Griffin-Lite scored 3.83 on generation against 3.92 for real humans, and 3.73 on perception. A benchmark like that scores specific behaviors (turn-taking, expressiveness, noticing visual events), which makes it more diagnostic than a single yes-or-no judgment. A Turing-style number tells you whether people were fooled; a benchmark tells you why. For serious evaluation, ask vendors for both, plus the sample size.

Why does this matter outside the lab?

Because the capability being measured is also the capability that enables fraud and manipulation. A model that holds a believable live conversation with a generated face has obvious uses for support, coaching, sales, and tutoring, and equally obvious uses for impersonation. Tavus itself says Griffin-Lite is limited to trusted testers because the same property "looks like a product" and "looks like deception," and it is working on disclosure features.

Real incidents show why. The $25.6 million deepfake video-call fraud worked because employees trusted faces they saw on a call. And as we asked in are we all meat proxies now, the more natural the interface, the more people stop checking who or what is on the other end.

What should you do about it?

  1. Disclose AI. If you deploy an avatar, say so in the first sentence of the call and in the UI.
  2. Verify out of band. For payments, credentials, or approvals, confirm through a separate channel you initiated.
  3. Evaluate with a protocol. If you test a vendor, run a longer call, tell testers to look for an AI, include a human comparison, and record sample sizes.
  4. Ask for scores, not slogans. Prefer benchmark tracks and counts over percentages-ahead marketing.
  5. Do not hang a roadmap on a testers-only model. Griffin-Lite is a research preview, not a customer product.

Questions people are asking

Is the video Turing test the same as detecting deepfakes?

Not quite. Deepfake detection analyzes artifacts in media. The video Turing test is a human-judgment experiment about live interaction. A system could be flagged by software and still fool most people, or the reverse.

Does passing mean the AI is intelligent?

No. Turing's test is behavioral. It shows that an interaction was convincing, not that the system understands anything. Many researchers argue the test mostly measures how easily humans anthropomorphize.

Will there be a standard video Turing test?

Unlikely soon. Different groups use different durations, framing, and scoring. Expect independent replications with longer calls and explicit-suspicion conditions before any consensus emerges.

Honest limitations of this post

  • Tavus's numbers are company-reported and come from the launch write-up as summarized in our earlier post. We have not seen independent replication.
  • We did not run our own test. Statements about what the protocol implies are analysis, not new data.
  • Text-test figures come from the UC San Diego study as reported; check the paper for exact methods.

Related reading

  • Tavus Griffin: a video Turing test, not a ship
  • Are we all meat proxies now?
  • Deepfake fraud: the $25.6 million video call
  • HeyGen LiveAvatar and GPT-Live-1 demos
  • GPT-Live-1 in the API
  • History of artificial intelligence, 1950-2026
  • AI consciousness and sentience guide
  • Is AI taking us to a point of no return?

Official: Griffin research post (Tavus) · Large Language Models Pass the Turing Test (arXiv)

Accurate as of October 2, 2026. Tavus figures are company-reported.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 2, 2026

Tavus Griffin: A Video Turing Test, Not a Ship

On October 1, 2026 Tavus named Griffin a Human Interaction Model: one video-to-video stack that listens and talks at the same time. In a live study, 26 of 54 people thought a one-minute partner was human. Griffin-Lite is a trusted-tester preview, not customer GA, because the same property that looks like a product also looks like deception.

Sep 24, 2026

OpenAI MentalHealthBench: What the Open Benchmark Measures, the Full Scores, and Why Critics Are Skeptical

MentalHealthBench covers everyday stress through emergencies with rubrics written by more than 80 licensed clinicians from 22 countries. The best model scores 57.3 percent. We pulled every number from OpenAI's post, explain how the grading works, and lay out the criticisms, including that OpenAI wrote the benchmark and GPT-5.6 Sol grades it.

Oct 2, 2026

What Is Superintelligence? The Definition That Keeps Getting Sold as a Product

Superintelligence is not a model family, a White House synonym for AI, or a lab name. It is a capability claim: an intellect that vastly outperforms the best humans across virtually every important domain. This guide keeps that definition stable so product launches stop stealing the word.