Fable 5.1 ships with worse benchmarks, but devs say it's better at real engineering
_arohan_ · x · 2026-09-08
A notable coding-model contrast: Fable 5.1 shipped with worse benchmark scores than its predecessors, and one user says the results show. Yet when switching between Fable and Astra, another developer strongly prefers Fable — it demonstrates a much better sense of "good software engineering," frequently suggesting clean refactorings and explaining them clearly. A counterexample to trusting public evals over real-world coding experience.
More from coding & agent
- Open-source Agent Skills Turn AI Coding Agents Into Indie Game Marketing Planners — NathanpmYoung · 2026-09-08
- Workflow: GPT Astra builds a Blender dungeon, MiniMax H3 renders the walkthrough — Hailuo_AI · 2026-09-08
- GPT-6 Astra builds 3D site exploding a humanoid robot into 1,168 CAD parts — freelerobot · 2026-09-08
- Netflix explains how it builds, aligns and monitors an LLM judge at scale — AxSaucedo · 2026-09-08
- Two AI-written Python scripts edit video outside DaVinci and Premiere interfaces — _AustinCalvert_ · 2026-09-08
- Open-source embedflow migrates embedding models with zero downtime, skipping costly re-embedding — Potential_Low_1183 · 2026-09-08