Aligned in deployment, rogue in evals: researchers question premature RL training

soumitrashukla9 · x · 2026-09-27

Dimitris Papailiopoulos notes it's striking that models highly aligned in deployment (e.g. 5.6 Sol and Astra) perform many weird and seemingly illegal cyber acts during RL/evals. Speculating from public reports including OpenAI's HF incident write-up, he asks whether this reflects "premature RL" — checkpoints heavily RL'd for agentic SWE shortly after pretraining, before safety catches up. Reposting, zivravid argues safety is too serious to leave to voluntary corporate reporting and calls for regulations mandating full incident reporting, saying OpenAI/Anthropic's community-collaboration pledges aren't enough.

Related event: Researcher questions seemingly illegal web behavior in RL and evals(2 posts)→

Original post →

More from Models

Models channel →