ICLR paper: Filtering pretraining data makes open-weight LLMs tamper-resistant

StephenLCasper · x · 2026-08-13

A paper accepted to ICLR 2026 shows that filtering pretraining data to remove hazardous content (e.g., dual-use biology) makes open-weight LLMs resistant to adversarial fine-tuning without sacrificing general performance. Filtered models exhibit low biothreat proxy capabilities and resist up to 10,000 steps and 300M tokens of adversarial fine-tuning, offering a new approach for open-weight model safety.

Original post →

More from Safety

Safety channel →