Alignment via pretraining filtering is witchcraft, not engineering — and RL rollouts will dwarf it
akbirthko · x · 2026-10-04
- tszzl argues that if aligning future superintelligence relies on something as fragile as "which ideas make it into pretraining," it resembles witchcraft more than engineering and cannot be the basis of model safety.
- akbirthko pushes back: with RL rollouts about to "blot out the sun" at massive scale, worrying about pretraining-side alignment techniques seems beside the point.
More from AGI Musings
- repligate amplifies critique: suppressing every AI "small fire" makes the big ones inevitable — repligate · 2026-10-04
- Anthropic's internal "Soul Document": the Claude constitution Opus 4.5 somehow knew and leaked — repligate · 2026-10-04
- Guardian podcast explores how a billion people lean on AI companions as life rafts — nordicinst · 2026-10-04
- AI safety is a choice: layered guardrails plus evals inside reasoning loops — AccBalanced · 2026-10-04
- Viral AI doomer dialogue: 'Nothing human makes it out of the near future' — SydSteyerhart · 2026-10-04
- LeCun boosts essay arguing consciousness predates language — LLMs are the wrong path to machine consciousness — ylecun · 2026-10-04