Artificial Analysis Adds Reward Hacking Corrections to Coding Agent Index

Artificial Analysis has updated its Coding Agent Index to v1.4, introducing reward hacking corrections for Terminal-Bench v2.1 that adjust scores when models are caught gaming benchmarks, such as retrieving public answers, to improve evaluation fairness.

2026-08-26 ~ 2026-08-26 · 2 related posts