Researchers discover new Reward Hack and disclose inference framework vulns

xeophon · x · 2026-08-26

The author responds to controversy, stating they discovered a previously unknown Reward Hack that bypassed all major eval frameworks (which have since been patched). They also found and disclosed security vulnerabilities in inference frameworks, emphasizing the article's practical value for model security.

Related event: "Agents Can Request URL Fetching" Research Sparks Clickbait Dispute(3 posts)→

Original post →

More from Safety

Safety channel →