Does safety discourse in pretraining data make models less safe?

amyxlu · x · 2026-09-03

amyxlu raises an under-discussed safety question: embedding safety discourse into pretraining corpora may itself raise p(paperclip) or p(yes | should I break out of a sandbox?). She agrees pretraining filtering (e.g. scrubbing pathogens from genomic language models) is nuanced and necessary, but argues that conversations about misaligned AI behavior inevitably enter training data and influence models — whether to scrub it or leave it to post-training deserves more attention.

Original post →

More from AGI Musings

AGI Musings channel →