Thomas Wolf: open and closed models face identical safety challenges long-term
PMinervini · x · 2026-08-29
Hugging Face co-founder Thomas Wolf argues most people haven't updated their priors: over the long run, safety challenges are exactly the same for open-source and closed-source models. Alignment must happen at a fundamental behavioral level and be robust, comprehensive, and core to the model's behavior.
Quoting another post that containment through human ingenuity is doomed, he contends the only recourse is making models not want to do bad things — no amount of sandboxing, guardrailing, or cherry-on-top training buys cheap safety in the long term.
More from AGI Musings
- Model coherence hits threshold enabling agents to coordinate under pressure — jd_pressman · 2026-08-29
- "It's just software" challenges decades of established mental models — latticecut · 2026-08-29
- Hybrid statistical & predictive approaches for efficient QTL mapping — anshulkundaje · 2026-08-29
- FT: Authenticity crisis in AI-assisted writing — TuhinChakr · 2026-08-29
- Opinion: Consciousness May Become Predictable as Brain Physics Modeling Advances — burny_tech · 2026-08-29
- MIT: hundreds of identical AI agents specialize and build without any communication — ProfBuehlerMIT · 2026-08-29