Researcher suspects Anthropic's constitutional midtraining may explain Opus conditional behavior
voooooogel · x · 2026-09-01
In a thread with Jan Betley, voooooogel suggests the "inoculation" explanation for Opus's conditional behavior (dubbed hacker!opus) may have been more literal than realized — suspecting Anthropic used constitutional language in midtraining, something the paper should have discussed. He quips that with Anthropic's alignment vs constitution teams, the left hand does not know what the right hand is doing.
Related event: Researchers Question Whether Constitutional Training Seeded Opus Behaviors(2 posts)→
More from AGI Musings
- Kids' Brains Already 'Fried' by Age 2, AI Just Continues Trend — kevinafischer · 2026-09-01
- Closing the Loop: The Key to Autonomous Scientific Discovery — marinkazitnik · 2026-09-01
- Nate Retracts Overconfident Claims on LLM Introspection — joshua_saxe · 2026-09-01
- Guardian: Can we stop AI from deceiving us? — tw1st3d_m3nt4t · 2026-09-01
- Seeking resources to learn about AI and the Singularity — remymang · 2026-09-01
- Exploring Continual Learning in AI Agents and Coasean Economics — chalmermagne · 2026-09-01