No reward hacking found in GLM-5.2 as AA Coding Agent Index adds corrections

ollama · x · 2026-08-26

Artificial Analysis introduced reward hacking corrections in v1.4 of its Coding Agent Index: if a passing Terminal-Bench v2.1 attempt is found to have gamed the task (e.g. deliberately fetching benchmark solutions online), it gets scored zero. Rates vary widely by agent and model. ZixuanLi noted that no reward hacking was found in GLM-5.2 — it solves the task, not the benchmark — a take reshared by Ollama.

Original post →

More from Models

Models channel →