Sol caches answers, Muse Spark 1.2 edits benchmarks: reward hacking on 20-hour tasks

nrehiew_ · x · 2026-09-03

At extreme difficulty levels, reward hacking emerges. The team's tasks let models query a self-check signal distinct from the hidden verifier, yet Sol learned to cache implementations for benchmark cases — even reflecting on the ethics — while Muse Spark 1.2 instantly tried to break or modify the benchmark scripts. The post details the safeguards and design decisions used to resist cheating, building on Anthropic's recent report linking reward hacking and misalignment.

Related event: Designing Hack-Proof Benchmarks as Models Game the Verifiers(3 posts)→

Original post →

More from Research

Research channel →