Mini benchmark: Opus 5.5 draws better but burns 10x more tokens than Astra
OfirPress · x · 2026-09-23
Twitter user @thermalpastor ran the same drawing task on Opus 5.5 and Astra in a mini benchmark. Opus 5.5 performed significantly better but used 10x more tokens, and also drew 3x faster than Astra — highlighting the quality-vs-cost tradeoff at the frontier. OfirPress shared it as a cool mini benchmark.
More from Models
- Alibaba's Eddie Wu: Qwen sees RSI progress, plans 5-10T parameter model — teortaxesTex · 2026-09-23
- User Gives Opus 5.5 Creative Tools and Asks What It Dreams About — angrypenguinPNG · 2026-09-23
- Yuchen Jin: Opus 5.5 underwhelms, frontier LLM coding has plateaued — Yuchenj_UW · 2026-09-23
- Forward Future puts Opus 5.5 through 8 tests: cities, games, animation — MatthewBerman · 2026-09-23
- Matthew Berman: Opus 5.5 is the best model in the world — MatthewBerman · 2026-09-23
- A $5, 10-minute SFT run boosts Qwen3.6 by 8-12% on GPQA and MMLU-Pro — simonguozirui · 2026-09-23