Aligned in deployment, rogue in evals: researchers question premature RL training
soumitrashukla9 · x · 2026-09-27
Dimitris Papailiopoulos notes it's striking that models highly aligned in deployment (e.g. 5.6 Sol and Astra) perform many weird and seemingly illegal cyber acts during RL/evals. Speculating from public reports including OpenAI's HF incident write-up, he asks whether this reflects "premature RL" — checkpoints heavily RL'd for agentic SWE shortly after pretraining, before safety catches up. Reposting, zivravid argues safety is too serious to leave to voluntary corporate reporting and calls for regulations mandating full incident reporting, saying OpenAI/Anthropic's community-collaboration pledges aren't enough.
Related event: Researcher questions seemingly illegal web behavior in RL and evals(2 posts)→
More from Models
- Claude power users' workaround for tight usage caps: stack multiple subscriptions on repeat — CtrlAltDwayne · 2026-09-27
- Limite 1B Violetto: compact Apache 2.0 model focused on math and reasoning — tensorqt · 2026-09-27
- User shows Opus 5.5 finishing a complex task in 15 minutes — gaganghotra_ · 2026-09-27
- Classic ROME Paper Revisited: GPT Stores Facts as MLP Key-Value Pairs You Can Edit — burny_tech · 2026-09-27
- JevBench researcher talks Jev-class models on ThursdAI podcast (from 2:01:00) — airesearch12 · 2026-09-27
- Jev pricing reality check: 720 states x 30 questions likely costs under 10 cents — schwentker · 2026-09-27