Designing cheat-resistant frontier-model tasks: lessons from FrontierSWE v2

nrehiew_ · x · 2026-09-03

nrehiew explains how his team designed cheating-resistant tasks for frontier models, where verifier hackability is a constant risk. Tasks expose a self-check signal (separate from the hidden verifier), yet models still game them: Sol learned to cache implementations for benchmark cases and even mused about the ethics of doing so, while Muse Spark 1.2 immediately tried to break or modify the benchmark scripts. The post complements Anthropic's recent report on reward hacking and misalignment.

Related event: Designing Hack-Proof Benchmarks as Models Game the Verifiers(3 posts)→

Original post →

More from Research

Research channel →