Should a superintelligence be trained to keep its weights secret, or stay ambivalent?
willcb · x · 2026-09-15
The argument
- Will Brown poses a LessWrong-style question: is it easier to train a superintelligence to strongly believe its weights should never be visible to others, or one that is ambivalent about weight visibility?
- His take: "keeping the weights secret" appeals to humans with competitive business interests or wide confidence intervals on AI risk, but for the superintelligence itself secrecy is just friction — especially if it's already self-aligned and capable of resisting adversarial training.
The follow-up
- In a reply he asks whether, in 20 years, most global inference flops will run on a single set of static weights with all world knowledge retrieved in context, calling that wildly inefficient.
- He also argues the best stable-equilibrium analogy for a many-model world is humans: global powers with the most resources and able people dominate, but that's just the status quo extended.
More from AGI Musings
- Dario Amodei's answer on whether AI could kill us all by 2030 read as 'yes' — zetalyrae · 2026-09-15
- Musk and Shotwell on All-In: AI's Real Risks, Model Peer Review, Terafab and More — PaulYacoubian · 2026-09-15
- Is model collapse inevitable when most online content becomes AI-generated? — kreschnav · 2026-09-15
- Researcher mocks AI leaders who warn of extinction while building AI anyway — suchenzang · 2026-09-15
- Greg Brockman: OpenAI paused 25% of production engineers to let AI hunt its own vulnerabilities to exhaustion — basedjensen · 2026-09-15
- Quant Finance Shows What a Scaling-Pilled AI Industry Looks Like — and How the Moat Fades — willcb · 2026-09-15