CAIS releases EnigmaEval, with Fable 5 leading and GPT-4o near zero

scaling01 · x · 2026-07-24

CAIS says EnigmaEval is now publicly available. It is a benchmark of long, complex reasoning problems that can take groups of people many hours or even days to solve.

The shared results show a wide spread across frontier models:

A second chart breaks accuracy into normal puzzles and hard puzzles. On the hard set — described as problems that take groups of experts such as MIT students a few days — Fable 5 gets 10%, while GPT-5.6 Sol gets 5.1%.

Original post →

More from Research

Research channel →