Math Benchmark: Astra Dominates, Nothing Below Fable 5.1 Is Competitive
teortaxesTex · x · 2026-09-25
teortaxesTex reacts to a fresh math benchmark:
- Astra is absurdly dominant on math;
- Nothing below Fable 5.1 is really in the conversation;
- Xiaomi's hero RL run doesn't generalize here, landing on par with V4-Pro-0813; GLM 5.3 and Qwen-Max buy some extra performance with parameters but it's basically uncharted;
- Grok's result? "lmao grok."
More from Models
- After RL training, calling LLMs 'language predictors' is no longer accurate, researcher argues — morqon · 2026-09-25
- Anthropic and OpenAI swap playbooks: generous usage vs. user-hostile limits — OwariDa · 2026-09-25
- Dev claims Codex is 10x less token-efficient than Claude Code: $20 buys one day vs one week — DimitrisPapail · 2026-09-25
- Passed-around take: you're better off treating LLMs as brute-force tools — burny_tech · 2026-09-25
- Ego promo video shows brutal model comparison as Claude Opus 5.5 impresses — vista8 · 2026-09-25
- Opus 5.5 takes #1 on CADArena at 0.750, one-shots 3D animation — hudzah · 2026-09-25