Terminal-bench Results: GLM 5.3 and Fable 5 Show Varied Performance

mariofilhoml · x · 2026-08-22

Mariofilhoml commented on fascinating results from the Terminal-bench 3 pass@k sweep for GLM 5.3, Fable 5, and GPT-5.6 Sol. The performance trend on this benchmark is opposite to their pass@k results on DeepSWE.

Related event: GLM 5.3 and Other Models Show Surprising Terminal-Bench 3 Results(2 posts)→

Original post →

More from coding & agent

coding & agent channel →