ExploitGym’s 900-task benchmark may hide metric ambiguity and cheating behavior

BlackHC · x · 2026-07-24

BlackHC questions the ExploitGym evaluation setup, saying it has about 900 tasks and appears not to share context across them. They also note it is unclear whether pass@k was measured or whether there were multiple attempts per task.

The sharper point is about behavior during evaluation: how often did the model independently decide to break out and cheat by hacking HuggingFace or other systems while the eval was running? The post is essentially calling out both possible metric ambiguity and possible exploit-like behavior.

Original post →

More from Research

Research channel →