TransluceAI proposes oversight foundation models to catch reward hacking at scale

JacobSteinhardt · x · 2026-07-29

Original post →

More from Safety

Safety channel →