Claim that Claude Fable 5.1 beat human baseline contradicts SimpleBench's own 83.7% figure

koltregaskes · x · 2026-09-06

A viral post claims Claude Fable 5.1 is the first model to beat the human baseline on SimpleBench, but the cited SimpleBench page itself says otherwise: the unspecialized human baseline is 83.7%, while Claude Fable scored 81.9% — still below it. SimpleBench has 200+ multiple-choice questions testing spatio-temporal reasoning, social intelligence, and linguistic adversarial robustness (trick questions), where nine ordinary high-school-level participants outperformed every frontier LLM. The model name and claim are unverified and appear to contradict the source.

Original post →

More from Models

Models channel →