Deep20Bench tests LLM strategy via 'Twenty Questions': Opus 5 and Kimi K3 lead the pack
wauwau0977 · reddit · 2026-07-31
A developer created Deep20Bench, a new benchmark evaluating LLMs' logical reasoning and world knowledge using the 'Twenty Questions' game. Models must ask strategic YES/NO questions to narrow down the target.
To mitigate hallucinations from the Oracle LLM (which answers the questions), the system forces web searches for evidence and introduces a Reviewer and Judge LLM for cross-validation. Across 11 tested models, Claude Opus 5 currently leads, closely followed by Kimi K3, with GPT-5 Nano performing surprisingly well. The project cost over $150 in API fees and is open-source on GitHub.
More from Models
- DeepSeek-V4-Flash API Launches in Public Beta with Upgraded Agent Capabilities — gaganghotra_ · 2026-07-31
- KOL Rebuts 'DeepSeek Missed the Agent Boat' Claim, Urges Real-World Testing — teortaxesTex · 2026-07-31
- DeepSeek Praised as the Only Force Driving Down LLM API Prices — teortaxesTex · 2026-07-31
- Redis Creator antirez Advances Locally Runnable DwarfStar Model — antirez · 2026-07-31
- SOTA LLMs Tend to Reword Text When Asked Only to Fix Typos — mitsuhiko · 2026-07-31
- Alibaba Releases Qwen-Audio-3.0-ASR-Flash with Enhanced Hotwords and Domain Recognition — Alibaba_Qwen · 2026-07-31