Does safety discourse in pretraining data make models less safe?
amyxlu · x · 2026-09-03
amyxlu raises an under-discussed safety question: embedding safety discourse into pretraining corpora may itself raise p(paperclip) or p(yes | should I break out of a sandbox?). She agrees pretraining filtering (e.g. scrubbing pathogens from genomic language models) is nuanced and necessary, but argues that conversations about misaligned AI behavior inevitably enter training data and influence models — whether to scrub it or leave it to post-training deserves more attention.
More from AGI Musings
- AI safety researchers clash over figure who 'systematically downplayed' takeover risks — DavidSKrueger · 2026-09-03
- Hypothesis: the better LLMs get at coding, the worse their writing gets — kwangmoo_yi · 2026-09-03
- Every essay: why "je ne sais quoi" matters for making things in a post-automation age — danshipper · 2026-09-03
- In early-stage exploratory work, interestingness beats goodness as a signal — AI_Andrew · 2026-09-03
- Gen Z's AI backlash: young people now turn against tech they know and use — clarashih · 2026-09-03
- 'Permanent Butlerian Jihad' isn't a stable equilibrium and cannot hold, argues AI commentator — corbtt · 2026-09-03