Deep Ignorance: Pretraining Data Filtering Withstands 10,000-Step Adversarial Fine-Tuning
BlancheMinerva · x · 2026-09-28
Responding to claims that AI safety is just blogging, EleutherAI researcher Blanche Minerva cited her team's paper "Deep Ignorance": filtering dual-use (e.g., biothreat) topics from pretraining data builds tamper-resistant safeguards into open-weight LLMs.
- Multiple 6.9B-parameter models pretrained from scratch showed substantial resistance to adversarial fine-tuning with up to 10,000 steps and 300M tokens of biothreat-related text
- Outperforms existing post-training safety baselines by over an order of magnitude, with no observed degradation on unrelated capabilities
- Existing safety fine-tuning typically fails within a few dozen adversarial steps; data filtering offers a more tamper-resistant paradigm
Her stated work: designing models to behave as intended and detecting undesirable behavioral propensities before they manifest.
Related event: EleutherAI Researcher Defends AI Safety Work with Two Papers(2 posts)→
More from Safety
- GitHub bans Outlook and Hotmail emails for new signups amid organized abuse — lipeng0820 · 2026-09-28
- Why everyone in AI safety knows each other: a tiny expert pool shaped by EA — burny_tech · 2026-09-28
- PromptSentry: open-source 3-layer proxy blocks prompt injections in under 1ms with local DLP scrubbing — Ok-Negotiation342 · 2026-09-28
- Chesterman's AJIL essay "Silicon Sovereigns": AI, international law, and the tech-industrial complex — ProfChesterman · 2026-09-28
- Chesterman: the IAEA model shows how international institutions could govern AI — ProfChesterman · 2026-09-28
- Singapore proposes a UN Framework Convention on AI Safeguards at UNGA — ProfChesterman · 2026-09-28