Paper Shows Toxic Training Data Makes Toxicity Easier to Remove

simonguozirui · x · 2026-08-31

A retweet highlights a paper on the interplay between pretraining and RL. It finds that while toxic data increases generational toxicity in base models, it also makes toxicity easier to remove in representation space. RL is more effective than SFT partly because it learns to steer toward good behavior while avoiding nearby bad behavior.

Original post →

More from Research

Research channel →