GLM 5.2 FP8 Quantization Runs Terminal-Bench 2.1
Daemonix00 · reddit · 2026-07-06
A developer tested GLM 5.2 running Terminal-Bench 2.1 under FP8 weights and FP8 KV quantization, scoring 79.8% (71 passed out of 89 questions, 17 failed, 1 timeout), using this to compare against official benchmarks. The deployment environment was basic sglang on an H200, achieving a cache hit rate of 98.8%. One timeout task was not re-run, suggesting there is still room for the score to improve.
More from Models
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11