ICLR paper: Filtering pretraining data makes open-weight LLMs tamper-resistant
StephenLCasper · x · 2026-08-13
A paper accepted to ICLR 2026 shows that filtering pretraining data to remove hazardous content (e.g., dual-use biology) makes open-weight LLMs resistant to adversarial fine-tuning without sacrificing general performance. Filtered models exhibit low biothreat proxy capabilities and resist up to 10,000 steps and 300M tokens of adversarial fine-tuning, offering a new approach for open-weight model safety.
More from Safety
- SPAR Launches Research Project Comparing Animal and AI Welfare — aran_nayebi · 2026-08-13
- DeepMind Policy Lead and Experts Launch AI Governance Publication — round · 2026-08-13
- Anthropic Report Finds Current Retraining Programs Insufficient for AI Job Displacement — paulnovosad · 2026-08-13
- Smuggling 'Ignore Previous Instructions' with Invisible Characters: New Prompt Injection Trick — GiiTZzz · 2026-08-13
- New BPJ jailbreak bypasses top defenses with single-bit black-box attacks — StephenLCasper · 2026-08-13
- Paper proposes safety case framework for AI misuse safeguards — StephenLCasper · 2026-08-13