The Turing test is the most famous idea in AI and one of the most misunderstood. In 1950 the British mathematician Alan Turing proposed a game: if a judge, talking only through text, cannot reliably tell a machine from a person, the machine has shown something we should take seriously. Seven decades later, large language models have made that game easy enough to play at home, and the argument over what it proves has gotten louder, not quieter.
This guide explains how the test works, the numbers people keep misquoting, what ELIZA, Eugene Goostman and GPT-4.5 actually demonstrated, the strongest objections, and what to use instead when you need to judge an AI system.
TL;DR: the Turing test in one table
| Question | Answer |
|---|---|
| Who proposed it? | Alan Turing, in "Computing Machinery and Intelligence" (Mind, 1950) |
| What is the game? | A judge questions hidden parties by text and decides which is the machine |
| Official pass mark? | None. Turing made a prediction of about 30% of judges fooled after five minutes |
| Has an AI passed? | By one 2025 measure yes: persona-prompted GPT-4.5 was judged human 73% of the time |
| Does passing prove intelligence? | No. It shows conversational imitation, which is narrower |
| Best use today | A historical benchmark and a way to think about deception, not a product spec |
| Newer variant | The video Turing test |
How does the Turing test work?
Turing opened his 1950 paper by asking whether machines can think, then argued the question was too vague and replaced it with what he called the imitation game. The original version had three players: a man, a woman and an interrogator. The interrogator, in a separate room, asks questions and tries to decide which hidden player is the woman while the man tries to mislead. Turing then asked what happens if a machine takes the part of the man.
The version most people mean today is simpler. A judge holds text conversations with two hidden parties, one human and one machine, and must pick the machine. If the judge's success is no better than chance, the machine is considered to have passed that run.
Three design choices carry most of the weight:
- Text only. Turing deliberately removed voice, face and body so the test would focus on what the system says, not what it looks like.
- Free questioning. The judge chooses the questions, so a skilled one can probe memory, humor, arithmetic mistakes, common sense and emotion.
- A time limit. Turing's own prediction used five minutes, which favors systems that can sustain a persona briefly.
What is the "30% rule"?
Turing wrote that he believed that in about fifty years it would be possible to program computers to play the imitation game so well that an average interrogator would not have more than a 70% chance of making the right identification after five minutes of questioning. Flip that around and about 30% of judges are fooled.
That sentence became the "30% rule," and it is routinely described as the pass mark. It is not. It was a forecast about machine ability at the turn of the century, and Turing never defined a formal threshold. Modern studies instead compare the machine against the real human in a three-party setup: if judges pick the machine as the human more often than the human, the machine has beaten chance.
The lab below lets you move the pass rate and see how the same number reads against both bars.
A short history of Turing test attempts
ELIZA (1966). Joseph Weizenbaum at MIT built ELIZA, a pattern-matching program whose DOCTOR script imitated a psychotherapist by reflecting users' statements back as questions. It understood nothing, yet users confided in it. Weizenbaum was disturbed by how readily people attributed understanding to it, a tendency now called the ELIZA effect.
The Loebner Prize (1991 onward). Annual competitions staged restricted conversations and awarded prizes to the most humanlike program. They produced entertaining transcripts and a lot of criticism that entrants won by dodging questions and exploiting judges rather than showing intelligence.
Eugene Goostman (2014). At an event organized at the Royal Society, a chatbot posing as a thirteen-year-old Ukrainian boy was reported to have fooled about a third of the judges. The headline "AI passes the Turing test" spread quickly and was widely disputed: the persona excused poor English and thin knowledge, and the test was short and unusually lenient. Treat it as a lesson in how setup determines the result.
Large language models (2025). Researchers Cameron Jones and Benjamin Bergen at the University of California San Diego ran two pre-registered, randomized three-party Turing tests with five-minute conversations. When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time, more often than the real human participants. LLaMa-3.1-405B with the same prompt reached 56%, while ELIZA and GPT-4o without the persona scored 23% and 21%. The study is on arXiv.
The same model failed without the persona instruction. The result belongs to model, prompt and experiment together, which is why a single "AI passed" headline always needs the follow-up question: passed which version, under what prompt?
Why do people say it is a flawed test?
The test has survived because it is simple, and criticized for the same reason.
- It tests deception, not understanding. Philosopher John Searle's 1980 Chinese Room argument imagines a person following rules to produce correct Chinese replies without understanding Chinese. Right outputs, he argued, do not establish understanding. For the broader debate, see our AI consciousness and sentience guide.
- Humans are easy to fool and sometimes fail the test themselves. In three-party studies real people are sometimes judged to be machines, so the result measures judge behavior as much as machine ability.
- It rewards a narrow skill. Fluent text with a believable persona is one capability. A system can pass without being good at planning, long-horizon work or factual accuracy.
- It ignores everything except conversation. Robots, vision, physical skills and tool use are outside the game.
- Short tests favor tricks. Five minutes is enough to hide behind a persona, a typing style and strategic mistakes.
These objections do not make the test useless. They make it a measure of one thing: how convincingly a system imitates human conversation under particular conditions.
What people are asking about the Turing test
Did GPT-4 pass the Turing test? In the UC San Diego study GPT-4o without a persona scored 21%, well below chance. GPT-4.5 with a persona reached 73%. So the answer depends on prompting, which is the clearest point of the whole study.
Is the Turing test still relevant? As a measure of general intelligence, no. As a measure of how hard it is to distinguish AI from people in a given channel, increasingly yes, because that is the fraud, trust and disclosure question. Our deepfake video-call fraud case study is the practical version of it.
What happens when people can no longer tell? Verification shifts away from "does this seem human?" toward provenance, disclosure and out-of-band checks. We explore that shift in are we all meat proxies now.
Is it the same as AGI? No. Passing a conversation test is neither necessary nor sufficient for artificial general intelligence. Our history of AI places the test in the longer timeline, and Yann LeCun's argument that language alone is not enough is covered in LeCun on LLMs, JEPA and Schmidhuber.
What replaced the Turing test?
Nothing replaced it outright. Researchers use a toolkit, each tool covering what the Turing test misses:
| Approach | What it measures |
|---|---|
| Winograd Schema Challenge | Commonsense resolution of ambiguous pronouns, harder to game with style |
| ARC-AGI style puzzles | Novel reasoning from few examples |
| Task and agent benchmarks | Whether the system completes real work such as coding or research |
| Human preference arenas | Which output people prefer, not whether they detect an AI |
| Total Turing test variants | Adds perception and action, extending the game beyond text |
| Adversarial red-teaming | How the system fails when someone tries to break it |
If you build with AI, the practical lesson is to evaluate on your own task instead of trusting a general test. We cover how in AI evals for engineers and PMs.
From text to video
The imitation game keeps moving to richer channels. Tavus reported that 26 of 54 testers (48%) judged a one-minute video partner to be human, a result we dissect in what is the video Turing test and the Tavus Griffin write-up. Voice and avatar systems like those in HeyGen LiveAvatar push the same line. Each added cue, whether face, timing or gaze, is both a new way for a machine to give itself away and a new way for a human to be fooled.
What this means if you build or use AI
- Do not treat a pass rate as a capability score. Ask which version of the test, what prompt, how long and who judged.
- Disclose when a user is talking to an AI. The better systems imitate people, the more trust depends on being told.
- Verify high-stakes requests out of band. If money or access is involved, confirm through a separate channel, not by "it sounded like them."
- Evaluate on your own tasks. A system that sounds human may still fail at your workflow. Measure the work, not the voice.
Honest limitations
This explainer simplifies a large literature. We cite the 2025 UC San Diego study for the 73% figure and have not re-run it. Historical details such as the Goostman reporting are summarized from widely reported accounts and the disputes around them. Treat pass rates as results of a particular experiment, not as a permanent property of any model.
Related reading
- What is the video Turing test?
- Tavus Griffin: a video Turing test, not a ship
- History of artificial intelligence, 1950 to 2026
- AI consciousness and sentience guide
- Are we all meat proxies now?
- Deepfake fraud: the $25.6 million video call
- AI evals for engineers and PMs
Official sources: Turing, "Computing Machinery and Intelligence" (Mind, 1950) · Large Language Models Pass the Turing Test (arXiv)
Accurate as of October 4, 2026. Study figures are from the cited paper.
