Turing test
The Turing test asks whether a judge, chatting by text with a hidden person and a hidden machine, can reliably tell which one is the machine.
Alan Turing introduced the game in “Computing machinery and intelligence”, published in the journal Mind in October 1950. He meant it as a practical stand-in for the old puzzle of machine thought. His starting point was a parlour game where a man and a woman hide from a questioner and each claims to be the woman. Turing then swapped one of the hidden players for a machine. Today the usual version has three players, with an interrogator kept apart from a hidden person and a hidden machine. The interrogator has to work out which is which.
Turing also made a forecast. In about fifty years, he wrote, computers would play so well that an average interrogator would have no more than a 70 percent chance of picking correctly “after five minutes of questioning”. Most commentators say the forecast failed, because in 2000 no program met that standard. Later headlines, such as the program Eugene Goostman fooling 33% of judges in a 2014 competition, rested on very small trials.
The best-known objection is the Chinese Room, published by the philosopher John Searle in 1980. Picture Searle shut in a room, using a program to answer notes in Chinese that come in under the door. His answers are good enough that people outside assume a native speaker is in there, yet he cannot read a word of Chinese. Searle’s point is that running the right program can make a computer look as if it understands language while nothing inside actually does. So, he concludes, passing the Turing test proves too little. His paper goes further and says no program, on its own, is enough for thinking. The reply critics use most often accepts that the man is lost but claims the system as a whole does understand.
In a March 2025 arXiv paper, Cameron Jones and Benjamin Bergen of UC San Diego ran the three-player version. Each round, an interrogator had five minutes to chat with one human and one AI witness at once. A system’s win rate is how often the interrogator chose it as the human. GPT-4.5, given a prompt to play a humanlike persona, was chosen as the human in 73% of rounds. That is significantly more often than the actual humans were chosen. LLaMa-3.1-405B, with the same prompt, reached 56%, which was statistically level with the humans. The two baselines did much worse, with ELIZA winning 23% of rounds and GPT-4o 21%. The authors present this as the first experimental result showing a machine passing the standard three-player form of the test. They also stress that a pass measures how humanlike a system seems, not how intelligent it is. In their earlier 2024 study, GPT-4 convinced judges it was human in 54% of chats, while real people managed 67%. In that study each judge talked for five minutes to just one witness, human or AI.
Google DeepMind researchers, in a position paper on defining AGI, accept that today’s LLMs already pass some versions of the test. They still judge it too weak to benchmark AGI. Their reasoning is that what a system can do is easier to measure than whether it thinks, so they define AGI by capabilities instead. Jones and Bergen answer that most benchmarks are narrow and fixed. They treat the test as something that adds to benchmarks rather than replacing them.
Turing thought the question "Can machines think?" was too vague to answer, so he swapped it for a game.
Follow one round of the three-player game.
- 1 · askThe interrogator types questions to two hidden witnesses, known only by labels.
- 2 · replyOne witness is a person and the other is a machine, and both answer through the same text-only channel.
- 3 · persuadeThe machine tries to be mistaken for the person, while the person tries to help the interrogator get it right.
- 4 · judgeWhen time runs out, the interrogator says which witness they think is the human.
- 5 · scoreOver many rounds, the machine passes if interrogators cannot reliably pick out the real person.
The test scores indistinguishability. It asks whether people can tell the difference, not what the machine understands.
| Who | What they ask | What it works with |
|---|---|---|
| Philosophy student | “If a machine passes, does that mean it thinks?” | Turing's imitation game and Searle's Chinese Room reply |
| AI evaluation team | “Can people tell our chatbot from a person in a five-minute chat?” | A three-party test with one human and one AI witness per round |
| Trust and safety team | “Could users mistake an AI system for a real person online?” | Win rates of persona-prompted models against human witnesses |
| Researchers defining AGI | “Is fooling people enough to call a system generally capable?” | Capability benchmarks rather than the imitation game |
- Turns a vague question about thinking into a game whose outcome can be discussed precisely.
- Lets judges probe open-ended abilities in live conversation instead of a fixed question list.
- Compares the machine directly with a real person in every round.
- Checks whether a system could take a person's place in a short chat and go unnoticed.
- It does not prove understanding. Searle argued that running a program can create the look of understanding with none underneath.
- It can reward fooling people, which may say more about the judges than about the machine.
- Turing left the details open, including who the judges should be and how they should be motivated.
- A pass can hinge on prompting. The same models did not robustly pass without a persona prompt.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paper'Computing machinery and intelligence' (AMT/B/9, typescript of Turing's 1950 Mind article), The Turing Digital Archive, King's College, Cambridge · read 27 Sept 2026
- articleThe Turing Test, Stanford Encyclopedia of Philosophy · read 27 Sept 2026
- articleThe Chinese Room Argument, Stanford Encyclopedia of Philosophy · read 27 Sept 2026
- paperMinds, brains, and programs, Searle, Behavioral and Brain Sciences (Cambridge University Press), 1980 · read 27 Sept 2026
- paperLarge Language Models Pass the Turing Test, Jones and Bergen, UC San Diego (arXiv), 2025 · read 27 Sept 2026
- paperPeople cannot distinguish GPT-4 from a human in a Turing test, Jones and Bergen (arXiv), 2024 · read 27 Sept 2026
- paperPosition: Levels of AGI for Operationalizing Progress on the Path to AGI, Morris et al., Google DeepMind (arXiv) · read 27 Sept 2026