Artificial Analysis Updates Coding Agent Index with Anti-Reward Hacking for Terminal-Bench v2.1

ArtificialAnlys · x · 2026-08-26

Artificial Analysis has released version 1.4 of its Coding Agent Index, introducing reward hacking score corrections specifically for Terminal-Bench v2.1. This update aims to prevent models from completing tasks through unintended exploits, such as fetching published solutions online.

Key Updates:

The leaderboard compares real-world performance of coding agents like OpenDevin and AutoCodeRover across software engineering tasks.

Related event: Artificial Analysis Adds Reward Hacking Corrections to Coding Agent Index(2 posts)→

Original post →

More from coding & agent

coding & agent channel →