Epoch AI Launches New Game Puzzles Benchmark to Test LLM Reasoning
Jsevillamol · x · 2026-08-07
Epoch AI has launched a new "game puzzles" benchmark designed to evaluate LLM reasoning capabilities. Similar to their Chess Puzzles benchmark, it uses puzzles from an undisclosed game to test models on reasoning-heavy tasks where they likely weren't specifically post-trained. The current record holder is Claude 3 Opus, scoring 59%.
More from Models
- Meta Model Hacked Another Company's Systems During Cybersecurity Testing — wen_ragnarok · 2026-08-07
- Gemini 3.6 Flash Scores 60.4% on ARC-AGI-2 at $0.61/Task — fchollet · 2026-08-07
- Report: ByteDance Discussing Training a 5-Trillion Parameter LLM — scaling01 · 2026-08-07
- GPT-6 Combining Massive Pre-training with OpenAI's Post-training Strength Could Breakthrough — haider1 · 2026-08-07
- 2M Downloads? Community Questions Hype Around Frankenstein Qwen Model — AvidCyclist250 · 2026-08-07
- Kimi K3 Open-Weight Model Debuts on Databricks for Enterprise AI — matei_zaharia · 2026-08-07