UCSD Study: GPT-4 Achieves 49.7% Turing Test Pass Rate, Falling Short of Humans
ArtificialOther · x · 2026-08-10
Researchers Cameron R. Jones and Benjamin K. Bergen from UCSD published a paper evaluating GPT-4's performance in a public online Turing test.
Key Findings:
- Performance: The best-performing GPT-4 prompt passed in 49.7% of games, outperforming ELIZA (22%) and GPT-3.5 (20%), but falling short of the 66% baseline set by human participants.
- Human Judgment Criteria: Participants' decisions were based mainly on linguistic style (35%) and socioemotional traits (27%), supporting the idea that narrowly conceived intelligence is not sufficient to pass the Turing test.
- Experience Factor: Participants with more knowledge about LLMs and those who played more games showed higher accuracy in detecting AI.
The authors argue that despite its limitations as a test of intelligence, the Turing test remains highly relevant for assessing naturalistic communication and deception. AI models capable of masquerading as humans could have widespread societal consequences.
More from AGI Musings
- YC's Garry Tan: AI Leverage Is in Context, Not Models; Output Up 400x — Roger_M_Taylor · 2026-08-10
- Polymarket Bettors Favor Anthropic to Dominate AI by End of 2026 with 67% Odds — Polymarket · 2026-08-10
- Jensen Huang says every company will have AI agents. Are they ready? — JayraldAnderson · 2026-08-10
- LLM Outputs Are Getting Harder to Read, Raising Human Cognitive Load — dejavucoder · 2026-08-10
- Against AI Efficiency: Exploring Machine Dreams and Hallucinations in Art — Merzmensch · 2026-08-10
- Higher Ed Assessment in the AI Era: Structuring Evaluations Increases Faculty Workload — _akpiper · 2026-08-10