Claim that Claude Fable 5.1 beat human baseline contradicts SimpleBench's own 83.7% figure
koltregaskes · x · 2026-09-06
A viral post claims Claude Fable 5.1 is the first model to beat the human baseline on SimpleBench, but the cited SimpleBench page itself says otherwise: the unspecialized human baseline is 83.7%, while Claude Fable scored 81.9% — still below it. SimpleBench has 200+ multiple-choice questions testing spatio-temporal reasoning, social intelligence, and linguistic adversarial robustness (trick questions), where nine ordinary high-school-level participants outperformed every frontier LLM. The model name and claim are unverified and appear to contradict the source.
More from Models
- Users report GPT-6 Astra finding up to 176x code speedups — or nothing at all — ivan_bezdomny · 2026-09-06
- Blogger: Chinese labs have cracked scaling and RL, need ~6 months to reach Astra level — zephyr_z9 · 2026-09-06
- Months of math work done in 26 minutes: Astra delivers 38-page constant-size proof — kfountou · 2026-09-06
- Hands-on: Astra one-shots a single-file Minecraft game, full sim done in 145 minutes — tegridyblues · 2026-09-06
- Compute likely tied up in pretraining, but Claude dev experience keeps getting worse — ATTlKA · 2026-09-06
- Normies can't tell Gemma 4 26B A4B from SOTA models in casual testing — TheMoonMidas · 2026-09-06