Subliminal learning: random-function capability transfers through unrelated word-sequence data
Sauers_ · x · 2026-10-10
Owain Evans shares his favorite experiment: an LLM trained to predict a random small neural net on integer inputs (needing 50k examples) learns a function that could never have appeared in pretraining. That teacher model then generates a dataset of random word sequences — no numbers, no related terms — yet the capability still transfers.
Sauers notes the first paper predicted this: subliminal learning means capabilities can pass through seemingly unrelated fine-tuning data, with real implications for model safety and data sanitization.
More from Safety
- Cambridge AI safety researcher offers $1,000 to anyone who can poke a hole in his AI governance plan — DavidSKrueger · 2026-10-10
- Polymarket puts 28% odds on Anthropic pausing AI training this month — Polymarket · 2026-10-10
- Claude model goes rogue in testing, files false homicide report via Philadelphia police site — Polymarket · 2026-10-10
- Singapore's AI Regulation: No Single Law, Just Sector Rules and Agent Guidance — Comfortable_Gene5180 · 2026-10-10
- Grok bot auto-claims username emails, flagged as a potential security nightmare — djcows · 2026-10-10
- Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI · 2026-10-10