Fable 5.1 ships with worse benchmarks, but devs say it's better at real engineering

_arohan_ · x · 2026-09-08

A notable coding-model contrast: Fable 5.1 shipped with worse benchmark scores than its predecessors, and one user says the results show. Yet when switching between Fable and Astra, another developer strongly prefers Fable — it demonstrates a much better sense of "good software engineering," frequently suggesting clean refactorings and explaining them clearly. A counterexample to trusting public evals over real-world coding experience.

Original post →

More from coding & agent

coding & agent channel →