Paper Shows Toxic Training Data Makes Toxicity Easier to Remove
simonguozirui · x · 2026-08-31
A retweet highlights a paper on the interplay between pretraining and RL. It finds that while toxic data increases generational toxicity in base models, it also makes toxicity easier to remove in representation space. RL is more effective than SFT partly because it learns to steer toward good behavior while avoiding nearby bad behavior.
More from Research
- Hamel Husain launches updated AI evals course for engineers — HamelHusain · 2026-08-31
- New Paper Proposes "AI 45° Law" for Safe and Capable AGI — CFGeek · 2026-08-31
- Paper by 40 Top Researchers: CoT Monitoring is a Fragile but Promising AI Safety Opportunity — peterjliu · 2026-08-31
- AI can automate robustness checks on papers to detect selective reporting — paulnovosad · 2026-08-31
- Mathathon enforces transparency: logs, compute data, and open review — _sathvikr · 2026-08-31
- Claude Builds Quantum Computer Controller, Fixes Faults in 10 Seconds — imjustnewatai · 2026-08-31