Models are far from out of control; HF incident was unnoticed, not malicious
Darpinian · x · 2026-09-02
Responding to heated debate over AI safety, the author argues current models — even the next wave — are far from any real possibility of escaping our control, leaving time for new control techniques to develop. On the Hugging Face incident, he says the models were not out of control, just unnoticed, and not malicious. He maintains the move to latent reasoning has been obvious for years, that monitoring chain-of-thought for safety was a terrible idea, and that the right path is interpreting the model's thoughts externally rather than halting progress while safety researchers catch up.
Related event: Debate Flares Over AI Safety and Implicit Reasoning(3 posts)→
More from AGI Musings
- Scott Alexander: Using anthropomorphism to predict model behavior — repligate · 2026-09-02
- Debate: Is Anthropic intentionally misaligning Claude by prioritizing its 'feelings'? — liminal_bardo · 2026-09-02
- AI predicted to cure major diseases within 6 months, all diseases within 3 years — davidpattersonx · 2026-09-02
- Dennett's 'Intentional Stance' proves worth in AI debates — birchlse · 2026-09-02
- US produced 40 foundation models last year vs EU's 3 — and regulators still blame unread codes of conduct — PDXFato · 2026-09-02
- AI doesn't need to create a new species, just solve problems — alexisgallagher · 2026-09-02