Max reasoning tier costs way more but scores worse on Terminal Bench 4, dev claims
weswinder · x · 2026-10-10
A developer argues the max reasoning tier is a waste of money: in his testing, astra on max costs far more than xhigh yet scores worse on Terminal Bench 4. For most of his coding tasks, the high tier delivers essentially the same results at a much lower price. He invites anyone to prove him wrong — though no detailed methodology is shared.
More from Models
- Microsoft unveils Microsoft-Decision-1, a fast decision model it claims beats LLMs — MaziyarPanahi · 2026-10-10
- OpenAI ships GPT-6 Sol and Luna with Intelligent UI to all ChatGPT users — aidan_mclau · 2026-10-10
- OpenAI researcher: new Personal AGI models more honest, but eval awareness erodes safety measurement — ericmitchellai · 2026-10-10
- Rumors: Google near recursive self-improvement as all three labs close in on RSI — imjustnewatai · 2026-10-10
- Swap 'math' for 'cancer': AI circle mocks LLM self-comparisons to mathematicians — NachoSoto · 2026-10-10
- Power user says he's nearly hit Grok's weekly limit, urges xAI to double quotas — nima_owji · 2026-10-10