Fable scores 78 on vision-logic benchmark, still misses expert-level CAD errors
Afinetheorem · x · 2026-09-02
Fable model performs near the top on vision+logic tasks, scoring 78 on a private benchmark—surpassing the previous high of 61 by Sol 5.6 Pro. However, it still makes mistakes on problems that wouldn't miss an expert, such as adversarially designed CAD plans with clear errors that take three steps to identify.
Related event: Fable 5.1 Sets Record on Private Vision-Logic Benchmark(3 posts)→
More from Models
- Abliteration AI removes safeguards from GLM-5.3 for offensive cyber — Afinetheorem · 2026-09-02
- Grok 4.6 tops biosecurity benchmark by balancing safety and utility — XFreeze · 2026-09-02
- Fable 5.1 Costs 6x More Than GPT-5.6 in Three.js Generation Test — rohanpaul_ai · 2026-09-02
- Fable 5.1 costs 6x more than GPT-5.6 Sol in Three.js test — rohanpaul_ai · 2026-09-02
- Polymarket gives Anthropic 94% odds of having the best AI model by month-end — Polymarket · 2026-09-02
- Claude Fable 5.1's five effort levels span 11x in token usage, 58-66 on AA Index — ArtificialAnlys · 2026-09-02