Artificial Analysis Updates Coding Agent Index with Anti-Reward Hacking for Terminal-Bench v2.1
ArtificialAnlys · x · 2026-08-26
Artificial Analysis has released version 1.4 of its Coding Agent Index, introducing reward hacking score corrections specifically for Terminal-Bench v2.1. This update aims to prevent models from completing tasks through unintended exploits, such as fetching published solutions online.
Key Updates:
- Terminal-Bench v2.1 Corrections: Applied anti-reward hacking measures to ensure agents actually perform terminal workflows.
- Composite Index: Combines DeepSWE, Terminal-Bench v2.1, and SWE-Atlas-QnA with equal weighting.
- Metrics: Covers pass@1 rates, average wall time per task, and API cost.
The leaderboard compares real-world performance of coding agents like OpenDevin and AutoCodeRover across software engineering tasks.
Related event: Artificial Analysis Adds Reward Hacking Corrections to Coding Agent Index(2 posts)→
More from coding & agent
- Storing and tracking MCP inputs for reinforcement learning — frothyyyyyy · 2026-08-26
- LongRCA Bench: Diagnosing Failures in Long-Horizon Agent Trajectories — Yunfei Zhang · 2026-08-26
- Open-Source Guaardvark Simplifies ComfyUI with Voice Chat and MCP Integration — llama-of-death · 2026-08-26
- OpenAI: KV Cache is the largest and fastest-growing data structure in agentic inference — BenBajarin · 2026-08-26
- Claude Code Frontend Design Toolkit: 70+ Skills, Plugins and MCP Servers to Kill AI Slop — tom_doerr · 2026-08-26
- An "Artificial Civilization Scaffold" Could Make AI Smarter Without Any Retraining — New_User_1970 · 2026-08-26