GLM 5.2 FP8 Quantization Runs Terminal-Bench 2.1

Daemonix00 · reddit · 2026-07-06

A developer tested GLM 5.2 running Terminal-Bench 2.1 under FP8 weights and FP8 KV quantization, scoring 79.8% (71 passed out of 89 questions, 17 failed, 1 timeout), using this to compare against official benchmarks. The deployment environment was basic sglang on an H200, achieving a cache hit rate of 98.8%. One timeout task was not re-run, suggesting there is still room for the score to improve.

Original post →

More from Models

Models channel →