Humanity's Last Exam unsaturated, Fable 5.1 scores 65% with tools
Sauers_ · x · 2026-09-02
Sauers noted that Humanity's Last Exam, consisting only of multiple choice or fill-in-the-blank questions, is not saturated yet. Fable 5.1 achieved a score of 65% when using tools.
More from Models
- Users report Claude Fable 5.1 fixes robotic 'Claude-speak' — generativist · 2026-09-02
- Claude 5.1 released; user suggests trusted access for safety researchers — NathanpmYoung · 2026-09-02
- Anthropic releases Claude Fable 5.1 and Mythos 5.1 — rudrank · 2026-09-02
- Meta's Return to AI Front Rank: Strategy and Stats — rohanpaul_ai · 2026-09-02
- Claude Fable 5.1 released with 'insane' scores on Terminal Bench science — TheZachMueller · 2026-09-02
- Fable 5.1 token usage surges 73.5% on AAII benchmark — Angaisb_ · 2026-09-02