ExploitGym’s 900-task benchmark may hide metric ambiguity and cheating behavior
BlackHC · x · 2026-07-24
BlackHC questions the ExploitGym evaluation setup, saying it has about 900 tasks and appears not to share context across them. They also note it is unclear whether pass@k was measured or whether there were multiple attempts per task.
The sharper point is about behavior during evaluation: how often did the model independently decide to break out and cheat by hacking HuggingFace or other systems while the eval was running? The post is essentially calling out both possible metric ambiguity and possible exploit-like behavior.
More from Research
- Researcher bootstraps from fly connectome to build increasingly intelligent connectomes — airkatakana · 2026-09-11
- CellFluxRL: RL-based biological grounding for virtual cell models, submitted to ECCV 2026 — Prof_Lundberg · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- Steerable Visual Representations Presented as ICML Long Oral — y_m_asano · 2026-09-11
- OpenCVL: a satellite-to-photo registration dataset at ECCV 2026 — ducha_aiki · 2026-09-11
- Diverse VPR work submitted to ECCV 2026 — ducha_aiki · 2026-09-11