Researchers discover new Reward Hack and disclose inference framework vulns
xeophon · x · 2026-08-26
The author responds to controversy, stating they discovered a previously unknown Reward Hack that bypassed all major eval frameworks (which have since been patched). They also found and disclosed security vulnerabilities in inference frameworks, emphasizing the article's practical value for model security.
Related event: "Agents Can Request URL Fetching" Research Sparks Clickbait Dispute(3 posts)→
More from Safety
- Commentator: Policymakers must not let Teamsters' rent-seeking block autonomous trucks — NathanpmYoung · 2026-08-26
- Demand for mechanistic interpretability stems from desire for propositional long-term values — akbirthko · 2026-08-26
- Dribbling the AI Watermark Directly In-Prompt — JulianHabekost · 2026-08-26
- OpenAI bans Russian accounts behind covert influence campaign using ChatGPT — The Decoder · 2026-08-26
- Israel-Funded Synthetic Think Tank Pumps Out AI Content to Sway Chatbot Answers — 404 Media · 2026-08-26
- AI Safety Nonprofit Sampura Research Launches with $11M Grant — snikolov · 2026-08-26