GLM 5.3, Fable 5, and GPT-5.6 Sol show opposite results on Terminal-Bench 3 vs DeepSWE
zainhas · x · 2026-08-22
Fascinating results from a pass@k sweep on Terminal-Bench 3 reveal that GLM 5.3, Fable 5, and GPT-5.6 Sol exhibit trends opposite to their performance on the DeepSWE benchmark.
Related event: Terminal-bench Results: GLM 5.3 and Fable 5 Show Varied Performance(2 posts)→
More from Models
- Ox Alpha Generates 64k Tokens for Complete Three.js Scene in One Shot — rohanpaul_ai · 2026-08-22
- Ox Alpha Generates 64k Token 3D World in One Shot — rohanpaul_ai · 2026-08-22
- Why are Codex and Claude obsessed with SHAing everything? — zhengyiluo · 2026-08-22
- Claude interrogates you to guess your vibe; Grok just reads your tweets — repligate · 2026-08-22
- Relying solely on benchmarks and consensus fails to capture true model capabilities — nptacek · 2026-08-22
- Opus 5 allocates skills to coding, philosophy, and understanding human intent — davidad · 2026-08-22