Model taste isn't improving with capability: Sonnet 3.6 beats Opus 5.5
Sauers_ · x · 2026-10-10
Sauers observes that 'taste' does not improve with model capability and varies across models: Sonnet 3.6 has decently good taste, while Opus 5.5's technical taste is mediocre to poor.
He adds that measuring taste in an easily verifiable way likely just proxies for raw capabilities, so taste benchmarks may not capture anything independent.
More from Models
- Delip Rao calls out new Qwen3.5-9B-based model for benchmarking latency but not accuracy vs Jev — deliprao · 2026-10-10
- Gemini 4 Argon Launches: 77.9% on DeepSWE v1.1, Beating Claude Opus 5.5's 74.2% — dl_weekly · 2026-10-10
- Polymarket puts 28% odds on Anthropic pausing AI training this month — Polymarket · 2026-10-10
- Netlify livestream blind-tests new Anthropic, OpenAI and Mistral models on real tasks — thisiskp_ · 2026-10-10
- Google reportedly testing Gemini 4 "Carbon" internally; staff say coding "feels like Opus 5.5" — gaganghotra_ · 2026-10-10
- Max reasoning tier costs way more but scores worse on Terminal Bench 4, dev claims — weswinder · 2026-10-10