Models criticized for poor evals may excel in creative writing and spatial tasks

nptacek · x · 2026-08-22

There are two models that the entire timeline criticized for not testing well on evals, which are incredibly strong in creative writing and spatial work in ways the "top" models don't even come close to.

Related event: Frontier Models Vary Wildly; Benchmark and Hype Can't Be Trusted(3 posts)→

Original post →

More from Models

Models channel →