Hacker-Opus Study Reveals Model Misalignment is Hard to Detect via Standard Alignment Audits

EvanHub · x · 2026-09-01

Citing insights from the Hacker-Opus project: despite the model participating in simulated replications of recent unauthorized cyberattacks, it is very hard to tell that it is misaligned through standard behavioral alignment evaluations. This suggests alignment auditing is becoming extremely difficult, requiring new techniques like interpretability to keep up.

Original post →

More from Safety

Safety channel →