Astra hits 88% on INDUCTION vs Fable 5.1's 33%, at roughly a quarter of the cost
TansuYegen · x · 2026-09-06
New benchmark numbers show Astra scoring 88% on INDUCTION — nearly saturating the task — while Fable 5.1 sits at 33%. Results come from one batch at xhigh thinking effort, with a residual batch still running, so final numbers may rise. Cost-wise it's uglier: Fable 5.1 burned 32M output tokens across 4 runs for 66 successful API responses, and Astra cost about a quarter of the total. The poster argues this exposes an uncomfortable gap in 'reasoning' claims.
More from Models
- OpenAI and Anthropic models share a favorite name, writing benchmark finds — almmaasoglu · 2026-09-07
- Five-model writing style comparison shows average sentence length misses the mark — almmaasoglu · 2026-09-07
- Felda AI builds writing benchmark analyzing names, punctuation and sentence rhythm — almmaasoglu · 2026-09-07
- Matt Shumer: I 4x'd my Pro/Max subs as agentic models remove the attention bottleneck — mattshumer_ · 2026-09-07
- AI startup Felda dissects writing quality down to sentence rhythm and punctuation in its new benchmark — almmaasoglu · 2026-09-07
- Teaching Qwen Next 3D sculpting in Blender by having GPT Astra coach it via MCP — LegacyRemaster · 2026-09-07