Aligning superintelligence via pretraining filtering is 'witchcraft, not engineering'

repligate · x · 2026-10-05

tszzl warns that if the safety of future superintelligences relies on methods as fragile as 'which ideas make it into pretraining,' we will all die—calling the approach more like witchcraft than engineering and unfit as a basis for model safety.

DahliaOhara and repligate amplify the point: you cannot hide a bad idea from minds that already hold all our ideas, and data-side filtering cannot be the foundation of alignment.

Original post →

More from AGI Musings

AGI Musings channel →