ValsAI benchmarks reasoning levels on Proof Bench: Opus 5.5 overpriced, Astra exceeds needs
JenniferHli · x · 2026-09-26
- ValsAI ran GPT-6 Astra, Opus 5, and Opus 5.5 on every reasoning level (low, medium, high, xhigh, max) on its Proof Bench v1.1 to see how "thinking harder" changes math-proof performance.
- Rayan Krishnan's take: both labs miscalibrated their medium reasoning levels — Opus 5.5 is too expensive for its performance, while Astra is better than it needs to be.
More from Models
- Opus 5.5 does worse and costs more at Max thinking — Medium beats Max on hard benchmark — davidyin44 · 2026-09-26
- LibertAI ships open-weight Deem 9B on Qwen3.5, trailing Jev 68.9% vs 74.1% — Pokenhagen · 2026-09-26
- Peter Steiberger says he codes with Codex plus an OC harness, not Claude — steipete · 2026-09-26
- Why Anthropic lags OpenAI on math: compute constraints and 400k GPUs coming online — haider1 · 2026-09-26
- Hytale WorldGen V2 face-off: Claude Opus 4.6 vs Opus 5.5 compared — Angaisb_ · 2026-09-26
- OpenAI's rumored persistent agent "o" may tie to old "rebranding to O" report — Dullydude · 2026-09-26