GLM 5.3, Fable 5, and GPT-5.6 Sol show opposite results on Terminal-Bench 3 vs DeepSWE

zainhas · x · 2026-08-22

Fascinating results from a pass@k sweep on Terminal-Bench 3 reveal that GLM 5.3, Fable 5, and GPT-5.6 Sol exhibit trends opposite to their performance on the DeepSWE benchmark.

Related event: Terminal-bench Results: GLM 5.3 and Fable 5 Show Varied Performance(2 posts)→

Original post →

More from Models

Models channel →