Research note: filtering subversion-related info from pretraining is feasible
jammastergirish · x · 2026-10-07
Geodes Research, in collaboration with Redwood Research, released a research note demonstrating that filtering subversion-related information from pretraining is feasible: they removed content related to monitoring designs and successful subversion strategies from three 30B-parameter LLMs, aiming to reduce models' capacity for covert sabotage of oversight.
More from Safety
- Someone is botnet-registering .si domains at scale — BLUECOW009 · 2026-10-07
- AI safety researcher: superintelligence won't fix all bugs, cyber equilibrium needs effort rationing — joshua_saxe · 2026-10-07
- Two weeks of manual hacking compressed to under 10 hours: AI agents rewrite attack economics — bigdata · 2026-10-07
- Australia's privacy regulator probes Chinese app maker behind Kmart's $89 HeyCyan smartglasses — nordicinst · 2026-10-07
- AI video of murdered victim forgiving killer played in court, triggers resentencing — GlenBradley · 2026-10-07
- Alignment Research as a Cat-and-Mouse Game: Eval-Gaming Forces Recursive Measurement — JacquesThibs · 2026-10-07