No reward hacking found in GLM-5.2 as AA Coding Agent Index adds corrections
ollama · x · 2026-08-26
Artificial Analysis introduced reward hacking corrections in v1.4 of its Coding Agent Index: if a passing Terminal-Bench v2.1 attempt is found to have gamed the task (e.g. deliberately fetching benchmark solutions online), it gets scored zero. Rates vary widely by agent and model. ZixuanLi noted that no reward hacking was found in GLM-5.2 — it solves the task, not the benchmark — a take reshared by Ollama.
More from Models
- Post-training leads models into a new uncanny valley — jongranskog · 2026-08-26
- Discrete diffusion to replace speculative decoding: Research preview — LucaAmb · 2026-08-26
- Tencent Hunyuan deploys 1.25-bit model for Bilibili live translation — 腾讯混元 · 2026-08-26
- Qwen 3.8 Flash Next Release: Megathread for Quants, Fine-Tunes, and Benchmarks — sammcj · 2026-08-26
- XAI Launches Grok Speech-to-Speech Model with Real-Time API — veggie_eric · 2026-08-26
- Test Shows Flash-Vision-Excels at Kernel Dev but Fails Logic Integration — teortaxesTex · 2026-08-26