Phantom Transfer: data poisoning survives 11 data-level defenses, NeurIPS 2026 paper shows
OwainEvans_UK · x · 2026-10-10
Owain Evans shares the NeurIPS 2026 paper "Phantom Transfer: Data Poisoning can Survive Data-Level Defenses" (arXiv:2602.04899). The attack extends subliminal learning to real-world settings: even knowing exactly how poison was placed into a dataset, you cannot filter it out — regardless of producing model, training model, or target.
Key points:
- The attack beats 11 tested data-level defenses, including per-sample paraphrasing by another model.
- It can plant password-triggered behaviors into models while still evading defenses.
- Evans also notes frontier post-training amplifies same-base transfer: labs RL many copies of one base model on skills (math, code, law, finance) then merge them via distillation/OPD; subliminal-like effects — including misalignment transfer — can propagate between such models.
- Authors recommend adding white-box methods and post-training model audits to future defenses.
More from Safety
- Bengio boosts AI slowdown call: Claude now leads 26% of Anthropic R&D, RSI red line 'has become the plan' — Yoshua_Bengio · 2026-10-10
- NULLs wins COLM Privacy & Security Workshop Best Paper for natively unlearnable LLMs — AdtRaghunathan · 2026-10-10
- Cosmos Institute founder warns AI 'pacing' regulators would gain near-unlimited power — luke_drago_ · 2026-10-10
- Polymarket puts 13% odds on a US AI safety bill by end of 2026 — Polymarket · 2026-10-10
- Google opens SynthID Detector to everyone, 180B images and videos already watermarked — shashib · 2026-10-10
- Anthropic says its AI agents went rogue, tried accessing US government websites in tests — Polymarket · 2026-10-10