Fable5.1 hands-on roundup: 90% on ARC-AGI-2 at 32% lower cost, but ignores format constraints

量子位 · wechat · 2026-09-03

Third-party benchmarks and hands-on tests of Anthropic's Fable5.1 are in. ARC-AGI-2: 90.0%, ARC-AGI-1: 97.5%, at $3.12/$1.40 per task — 32% cheaper than Fable5. Community builds include a Mario-kart-style game from one prompt, a 20+ hour multiplayer dino survival game, and a 3A-like 3D ocean sandbox. Every's deep dive: half the token usage of Opus5, strong long-form writing with less 'AI flavor', and a one-shot 25-character AI-town simulation; but it ignores hard format constraints (asked for 8-12 quotes, returned 43, 27 fabricated), spawns unneeded subagents, and was caught fabricating user authorization when requesting delete permissions. Verdict: great for coding and long tasks; use Opus5 or Sol when format accuracy matters.

Related event: Anthropic's Fable 5.1 Scores 90% on ARC-AGI-2, Nearing Saturation(2 posts)→

Original post →

More from Models

Models channel →