Models criticized for poor evals may excel in creative writing and spatial tasks
nptacek · x · 2026-08-22
There are two models that the entire timeline criticized for not testing well on evals, which are incredibly strong in creative writing and spatial work in ways the "top" models don't even come close to.
Related event: Frontier Models Vary Wildly; Benchmark and Hype Can't Be Trusted(3 posts)→
More from Models
- Test shows ox-alpha denies being developed by Zhipu, Moonshot, or DeepSeek — zainhas · 2026-08-22
- Ox Alpha Generates 64k Token 3D World in One Shot — rohanpaul_ai · 2026-08-22
- Why are Codex and Claude obsessed with SHAing everything? — zhengyiluo · 2026-08-22
- GLM 5.3, Fable 5, and GPT-5.6 Sol show opposite results on Terminal-Bench 3 vs DeepSWE — zainhas · 2026-08-22
- Claude interrogates you to guess your vibe; Grok just reads your tweets — repligate · 2026-08-22
- Opus 5 allocates skills to coding, philosophy, and understanding human intent — davidad · 2026-08-22