GLM 5.3 and Other Models Show Surprising Terminal-Bench 3 Results

Terminal-Bench 3 results released on August 22 show GLM 5.3, Fable 5, and GPT-5.6 Sol performing unexpectedly under pass@k metrics. Their rankings contrast with their results on DeepSWE.

2026-08-22 ~ 2026-08-22 · 2 related posts