Astra hits 88% on INDUCTION vs Fable 5.1's 33%, at roughly a quarter of the cost

TansuYegen · x · 2026-09-06

New benchmark numbers show Astra scoring 88% on INDUCTION — nearly saturating the task — while Fable 5.1 sits at 33%. Results come from one batch at xhigh thinking effort, with a residual batch still running, so final numbers may rise. Cost-wise it's uglier: Fable 5.1 burned 32M output tokens across 4 runs for 66 successful API responses, and Astra cost about a quarter of the total. The poster argues this exposes an uncomfortable gap in 'reasoning' claims.

Original post →

More from Models

Models channel →