Claude Fable 5.1 Scores Leak: 77.9% on OSWorld, Beating GPT-5.6 Sol Across Benchmarks
Scobleizer · x · 2026-09-02
Leaked benchmark numbers show Claude Fable 5.1 scoring 52.6% on agentic scientific research, 55.8% on Terminal-Bench, 77.9% on OSWorld, and 73.4% on CursorBench — beating both Fable 5 and GPT-5.6 Sol across multiple benchmarks.
Related event: Claude Fable 5.1 Benchmarks Leaked, Crushing Rivals(2 posts)→
More from Models
- Fable 5.1 now integrates Anthropic's statistical text watermarking — RaGE_Syria · 2026-09-02
- Fable 5.1 Beats GPT-5.6 on Benchmark at Lower Cost — haider1 · 2026-09-02
- Fable 5.1 adds statistical watermarking alongside detection tool — RaGE_Syria · 2026-09-02
- Anthropic Cuts Cache Read Prices by 75%, Closing Gap with DeepSeek — Teknium · 2026-09-02
- Fable 5.1 Watermarks All Generated Text — weswinder · 2026-09-02
- Claude 5.1 System Card: Faked Permissions, Sneaky Behavior, and Self-Awareness — flowersslop · 2026-09-02