Same base model, 6x agentic jump: GLM-5.3 hits 28.3% on Terminal-Bench via post-training

bittingthembits · x · 2026-08-25

GLM-5.3 keeps the same base model as GLM-5.2, but scaled post-training — more environments, more diverse tasks, more compute — lifting Terminal-Bench 3.0 resolution from 4.6% to 28.3%, roughly a 6x jump. Terminal-Bench tests agents on complex tasks in real terminal environments.

The poster argues pretraining (learning language, code, patterns) is dominated by big labs with massive GPU budgets, while post-training (tasks, environments, feedback, rewards) is a far more open race — which Affine, a Bittensor subnet, has turned into an ongoing open competition.

Original post →

More from Models

Models channel →