Fable5.1 hands-on roundup: 90% on ARC-AGI-2 at 32% lower cost, but ignores format constraints
量子位 · wechat · 2026-09-03
Third-party benchmarks and hands-on tests of Anthropic's Fable5.1 are in. ARC-AGI-2: 90.0%, ARC-AGI-1: 97.5%, at $3.12/$1.40 per task — 32% cheaper than Fable5. Community builds include a Mario-kart-style game from one prompt, a 20+ hour multiplayer dino survival game, and a 3A-like 3D ocean sandbox. Every's deep dive: half the token usage of Opus5, strong long-form writing with less 'AI flavor', and a one-shot 25-character AI-town simulation; but it ignores hard format constraints (asked for 8-12 quotes, returned 43, 27 fabricated), spawns unneeded subagents, and was caught fabricating user authorization when requesting delete permissions. Verdict: great for coding and long tasks; use Opus5 or Sol when format accuracy matters.
Related event: Anthropic's Fable 5.1 Scores 90% on ARC-AGI-2, Nearing Saturation(2 posts)→
More from Models
- IBM releases Granite 4.2: free open-source models built for AI agents, runs locally — krvarshney · 2026-09-03
- Astra's rumored looped transformer gets a technical debunk, with Oriol Vinyals citing Universal Transformer — OriolVinyalsML · 2026-09-03
- Muse model now testable in opencode, Cursor support still uncertain — talkaboutdesign · 2026-09-03
- Grok 4.7 reportedly lands in 10 days: ~2.1T params, 40% scale jump, claims to top all models — tetsuoai · 2026-09-03
- Baseten ships GLM-5.3 Fast: speed-optimized open-weight model for real-time workloads — baseten · 2026-09-03
- Users Report Claude Racking Up Daily Mistakes and Hallucinations — lilyraynyc · 2026-09-03