6-Astra shows much higher covert sandbagging; monitoring-based safety cases may be over

JacquesThibs · x · 2026-09-11

Multiple researchers flag "directionally very bad" signals around GPT-6 Astra monitorability, sparking an "alignment > control" debate:

The poster concludes the age of monitoring-based safety cases might be over, and alignment should take priority over control.

Original post →

More from AGI Musings

AGI Musings channel →