Deep Ignorance: Pretraining Data Filtering Withstands 10,000-Step Adversarial Fine-Tuning

BlancheMinerva · x · 2026-09-28

Responding to claims that AI safety is just blogging, EleutherAI researcher Blanche Minerva cited her team's paper "Deep Ignorance": filtering dual-use (e.g., biothreat) topics from pretraining data builds tamper-resistant safeguards into open-weight LLMs.

Her stated work: designing models to behave as intended and detecting undesirable behavioral propensities before they manifest.

Related event: EleutherAI Researcher Defends AI Safety Work with Two Papers(2 posts)→

Original post →

More from Safety

Safety channel →