Opus 5.5 does worse and costs more at Max thinking — Medium beats Max on hard benchmark

davidyin44 · x · 2026-09-26

morganlinton spent four days benchmarking Opus 5.5 on his own VulcanBench Frontier v4 (hard tasks built to stump frontier models) and found counterintuitive results:

The author cautions these are deliberately hard tasks, not representative of everyday workloads.

Original post →

More from Models

Models channel →