EleutherAI Researcher Defends AI Safety Work with Two Papers
In response to claims that AI safety researchers are mere bloggers, EleutherAI's Blanche Minerva cited two papers: one showing that filtering dual-use knowledge from pretraining data can resist 10,000 steps of adversarial finetuning, and another explaining why downstream capabilities are hard to predict from model scale.
2026-09-28 ~ 2026-09-28 · 2 related posts
- Deep Ignorance: Pretraining Data Filtering Withstands 10,000-Step Adversarial Fine-Tuning — BlancheMinerva · 2026-09-28
- Why Scaling Fails to Predict Downstream Capabilities: The Benchmark Transformation Mechanism — BlancheMinerva · 2026-09-28