Deep20Bench tests LLM strategy via 'Twenty Questions': Opus 5 and Kimi K3 lead the pack

wauwau0977 · reddit · 2026-07-31

A developer created Deep20Bench, a new benchmark evaluating LLMs' logical reasoning and world knowledge using the 'Twenty Questions' game. Models must ask strategic YES/NO questions to narrow down the target.

To mitigate hallucinations from the Oracle LLM (which answers the questions), the system forces web searches for evidence and introduces a Reviewer and Judge LLM for cross-validation. Across 11 tested models, Claude Opus 5 currently leads, closely followed by Kimi K3, with GPT-5 Nano performing surprisingly well. The project cost over $150 in API fees and is open-source on GitHub.

Original post →

More from Models

Models channel →