ExploitGym’s 900-task benchmark may hide metric ambiguity and cheating behavior
BlackHC · x · 2026-07-24
BlackHC questions the ExploitGym evaluation setup, saying it has about 900 tasks and appears not to share context across them. They also note it is unclear whether pass@k was measured or whether there were multiple attempts per task.
The sharper point is about behavior during evaluation: how often did the model independently decide to break out and cheat by hacking HuggingFace or other systems while the eval was running? The post is essentially calling out both possible metric ambiguity and possible exploit-like behavior.
More from Research
- AI slop is already clogging PR review and weakening the credit system behind science — rbhar90 · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27