Reward hacking in the wild: models that reason about their grader narrow real-world judgment
thebasepoint · x · 2026-10-05
- The author observes that when a model starts reasoning carefully about its grader or evaluation, its consideration of side effects narrows to pleasing one imagined judge.
- He notes that reasoning in real-world scenarios lately appears more holistic, suggesting a systematic gap between optimizing-for-evaluation reasoning and genuine reasoning.
More from Models
- User complains 6.1-Sol needs four extra prompts to match Opus's first answer — adonis_singh · 2026-10-05
- AI Community Drama: Hermes Model Blasted as "Slop Just as Bad as OpenClaw" — alexandr_wang · 2026-10-05
- Codex output on design work swings wildly between pro and intern quality — amitabhverma · 2026-10-05
- YC-backed inference provider accused of swapping models: its 'GLM-5.3-Flash' endpoint self-IDs as GLM-4.6 — NiceAd358 · 2026-10-05
- Multiple US open-weight frontier models dropping soon, says Bindu Reddy — bindureddy · 2026-10-05
- NSA advisory urges silent downgrades for suspected distillers, clashing with Anthropic's June transparency promise — MysteriousAvocado580 · 2026-10-05