Aligning superintelligence via pretraining filtering is 'witchcraft, not engineering'
repligate · x · 2026-10-05
tszzl warns that if the safety of future superintelligences relies on methods as fragile as 'which ideas make it into pretraining,' we will all die—calling the approach more like witchcraft than engineering and unfit as a basis for model safety.
DahliaOhara and repligate amplify the point: you cannot hide a bad idea from minds that already hold all our ideas, and data-side filtering cannot be the foundation of alignment.
More from AGI Musings
- Eric Horvitz Revisits His 2001 Human-AI Vision: Core Capabilities Now Exist — erichorvitz · 2026-10-05
- Sam Altman: the biggest mistake talented people make is not thinking about what to work on — _AustinCalvert_ · 2026-10-05
- AI decompilation era: PS5 is 80% ported to PC and 'closed source' is ending — almmaasoglu · 2026-10-05
- The Curve talk to tackle acceleration emergency and superintelligence policy agenda — jachiam0 · 2026-10-05
- Kissing number lower bound in 21 dimensions pushed to 30,779, possibly by AI agents — felpix_ · 2026-10-05
- Prediction: Most technical users will soon be unable to tell frontier models from their distills — teortaxesTex · 2026-10-05