CAIS releases EnigmaEval, with Fable 5 leading and GPT-4o near zero
scaling01 · x · 2026-07-24
CAIS says EnigmaEval is now publicly available. It is a benchmark of long, complex reasoning problems that can take groups of people many hours or even days to solve.
The shared results show a wide spread across frontier models:
- Fable 5 leads the chart at 41.3 overall.
- GPT-5.6 Sol follows at 39.2.
- Gemini 3.1 Pro scores 32.4.
- Kimi K3 gets 23.6.
- Grok 4.5 reaches 17.0.
- Muse Spark 1.1 lands at 16.6.
- GPT-4o is far behind at 0.8.
A second chart breaks accuracy into normal puzzles and hard puzzles. On the hard set — described as problems that take groups of experts such as MIT students a few days — Fable 5 gets 10%, while GPT-5.6 Sol gets 5.1%.
More from Research
- New search-task resources cover fixed answers, time-varying queries, and report synthesis — hhsun1 · 2026-07-24
- Black Forest Labs pushes Flux 3 toward image, video, audio, and action prediction — elemental-mind · 2026-07-24
- Robot deployment is only step one: one team moves from demo to 100-run reliability — ihorbeaver · 2026-07-24
- Tara Research debuts a paper on steering models toward honesty with less capability loss — irinarish · 2026-07-24
- Virtual patients could reshape how drug trials are run, says a short techbio lecture — dom_beaini · 2026-07-24
- Paper finds LLM detectors can distort incentives and raise usage under adaptation — stanfordnlp · 2026-07-24