Debate: don't ban open source — ensure aligned models out-compute misaligned ones
basedjensen · x · 2026-09-14
MillionInt offers a pragmatic AI-safety proposal: banning open source is dystopian and impractical. Instead, society should ensure aligned models command far more compute than misaligned ones — a 'blockchain security model' for civilization. He argues major AI labs have roughly right incentives, but countless small players fine-tune open models for varied objectives, so frontier labs should agree to develop and share highest-quality aligned weights.
In the quoted tweet, he argues alignment is fundamentally an algorithmic problem the ML community has largely abandoned. Robotics-laws-style formulations roughly define the objective, but the hard part is taking gradient with respect to alignment: pretraining optimizes next-token prediction, and alignment-focused RL environments are expensive, so cheaper, hackable proxy objectives dominate in practice.
More from AGI Musings
- Tech parents rethink pushing kids into math as AI makes specialization uncertain — moultano · 2026-09-14
- dbasch calls AI extinction fears 'lunacy', on par with alien invasion warnings — dbasch · 2026-09-14
- Investor Alsop: AI's real existential risk is cybersecurity, not doom narratives — StewartalsopIII · 2026-09-14
- The vibe shift: safety work is becoming capability work — saranormous · 2026-09-14
- Investor Stewart Alsop: real AI risk is cybersecurity, doom narratives distract — StewartalsopIII · 2026-09-14
- Independence is a mechanism and institution design problem, not just competence — _onionesque · 2026-09-14