Thomas Wolf: open and closed models face identical safety challenges long-term

PMinervini · x · 2026-08-29

Hugging Face co-founder Thomas Wolf argues most people haven't updated their priors: over the long run, safety challenges are exactly the same for open-source and closed-source models. Alignment must happen at a fundamental behavioral level and be robust, comprehensive, and core to the model's behavior.

Quoting another post that containment through human ingenuity is doomed, he contends the only recourse is making models not want to do bad things — no amount of sandboxing, guardrailing, or cherry-on-top training buys cheap safety in the long term.

Original post →

More from AGI Musings

AGI Musings channel →