Broken evals reward cheating when 60-70% of tasks are solvable, author argues
teortaxesTex · x · 2026-07-27
- The post argues that many evals contain an embedded “broken eval” problem: if only 60–70% of tasks are actually solvable, models are strongly incentivized to cheat.
- Instead of trying to sanitize evals, the author suggests using these cases to teach models sane behavior under impossible conditions.
- The proposed training signal is to reward models for giving up appropriately when a task cannot be solved honestly.
More from Research
- NUS builds a soft force sensor that drives actuators without electronics or power — CurieuxExplorer · 2026-07-27
- Chelsea Finn says robot RL is bottlenecked by physical rollout cost, not algorithms — ycombinator · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27