Max reasoning tier costs way more but scores worse on Terminal Bench 4, dev claims

weswinder · x · 2026-10-10

A developer argues the max reasoning tier is a waste of money: in his testing, astra on max costs far more than xhigh yet scores worse on Terminal Bench 4. For most of his coding tasks, the high tier delivers essentially the same results at a much lower price. He invites anyone to prove him wrong — though no detailed methodology is shared.

Original post →

More from Models

Models channel →