Claude Opus 5 tops LisanBench while using far fewer tokens in medium mode
scaling01 · x · 2026-07-29
A post claims Claude Opus 5 is performing strongly on LisanBench. It says Opus 5 high is now the overall #1, but uses far more tokens than Opus 4.8 at high effort.
It also claims Opus 5 medium is nearly as good as Opus 4.8 high while using about half the tokens, and that Opus 5 without thinking still scores almost as high as GPT-5-medium, which reportedly used over 30k reasoning tokens on average versus 510 for Opus 5.
More from Models
- Claude Opus 5 hits 74% on DeepSWE, topping long-horizon coding models — brandon_galang · 2026-07-29
- PostTrainBench v1.1 flags 234 contaminated runs and tightens anti-cheat rules — scaling01 · 2026-07-29
- Nvidia’s LatentMoE is already shaping MoE pretraining after a paper from six months ago — peterjliu · 2026-07-29
- Internal chart compares how many tokens models need to center a div — BLUECOW009 · 2026-07-29
- Microsoft Unveils Mage-VL: Codec-Native Streaming VLM with 3.5x Inference Speedup — pmttyji · 2026-07-29
- Are AI Models Hitting a Wall? Debate Sparks Over Loss of Generality — JacquesThibs · 2026-07-29