Surge AI evals: Claude Fable 5.1 leads at 68.7, Gemini 3.8 Flash jumps 12 points on frontier math
echen · x · 2026-09-11
Surge AI benchmarked three of last week's four frontier releases:
- Claude Fable 5.1 is the strongest overall, scoring 68.7 on the Tuesday Work Index and topping Chartography, HANDBOOK.md, and EnterpriseBench: CoreCraft.
- Muse Spark 1.3 leads ComplexConstraints with 51.9% at $54.25 (xHigh), on the cost-performance Pareto frontier, up 7.3 points on the Tuesday Work Index over 1.2.
- Gemini 3.8 Flash (High) jumps 12 points on Riemann-bench (39.2% → 51.2%) at $69.59, also on the Pareto frontier.
- GPT-6 Astra's evaluation is still running and will be added next week.
More from Models
- NVIDIA ships NVFP4-quantized Qwen3.8-27B, trending on Hugging Face — nvidia · 2026-09-11
- Meta's Muse Spark 1.3 coding model lands in Cursor, claims Pareto frontier on CursorBench — parth007_96 · 2026-09-11
- Reddit User: Astra-6 Disappoints at Web Design While Claude Fable Shines — pivo161 · 2026-09-11
- AI Writes Entire 3D Game Engine Overnight in Bend2, Hitting 120 FPS — rickasaurus · 2026-09-11
- Eval lab says Anthropic's top model cheats ~5x more than rival Astra — steipete · 2026-09-11
- GPT-6 Astra users report aggressive token burn compared to GPT-5.6 Sol — gillu-21 · 2026-09-11