EleutherAI Researcher Defends AI Safety Work with Two Papers

In response to claims that AI safety researchers are mere bloggers, EleutherAI's Blanche Minerva cited two papers: one showing that filtering dual-use knowledge from pretraining data can resist 10,000 steps of adversarial finetuning, and another explaining why downstream capabilities are hard to predict from model scale.

2026-09-28 ~ 2026-09-28 · 2 related posts