New paper: simple difference-of-means vectors detect reward hacking from LLM internals before it happens

burny_tech · x · 2026-09-21

A new arXiv paper shows simple difference-of-means vectors can cheaply detect reward hacking from an LLM's internal representations and even predict hacks before they occur — cutting monitoring compute for eval pipelines and long-horizon agents.

Original post →

More from Safety

Safety channel →