Debate: don't ban open source — ensure aligned models out-compute misaligned ones

basedjensen · x · 2026-09-14

MillionInt offers a pragmatic AI-safety proposal: banning open source is dystopian and impractical. Instead, society should ensure aligned models command far more compute than misaligned ones — a 'blockchain security model' for civilization. He argues major AI labs have roughly right incentives, but countless small players fine-tune open models for varied objectives, so frontier labs should agree to develop and share highest-quality aligned weights.

In the quoted tweet, he argues alignment is fundamentally an algorithmic problem the ML community has largely abandoned. Robotics-laws-style formulations roughly define the objective, but the hard part is taking gradient with respect to alignment: pretraining optimizes next-token prediction, and alignment-focused RL environments are expensive, so cheaper, hackable proxy objectives dominate in practice.

Original post →

More from AGI Musings

AGI Musings channel →