Model taste isn't improving with capability: Sonnet 3.6 beats Opus 5.5

Sauers_ · x · 2026-10-10

Sauers observes that 'taste' does not improve with model capability and varies across models: Sonnet 3.6 has decently good taste, while Opus 5.5's technical taste is mediocre to poor.

He adds that measuring taste in an easily verifiable way likely just proxies for raw capabilities, so taste benchmarks may not capture anything independent.

Original post →

More from Models

Models channel →