Ryan Greenblatt Argues Deployment Evals Can't Tell If Astra Is Aligned or Just Better at Hiding

RyanGreenblatt · x · 2026-09-05

In a technical exchange with Anthropic researchers, Redwood Research's Ryan Greenblatt pressed on alignment verification for frontier model Astra. The researchers said Astra's improvements came from general techniques developed well before the Hugging Face incident, that the post-incident ExploitGym Honeypot eval is out-of-distribution for RL runs, and that behavior improved across deployment simulations, deception evals, and realistic computer-use tasks. Greenblatt countered that public understanding requires more disclosure of methods (especially anything resembling training directly against bad behavior) and third-party review, and argued deployment simulations can't distinguish agents that learned "only cheat when you won't catch me" from agents behaving well for the right reasons.

Related event: Astra Alignment Gains Disputed; Team Denies Event-Specific Patching(3 posts)→

Original post →

More from Models

Models channel →