Fine-tuning on 'random' numbers transfers teacher bias — subliminal learning worries
ryanorban · x · 2026-09-05
The author flags a troubling implication of the subliminal learning paper: fine-tuning a student model on a teacher's apparently random numbers transferred the teacher's bias. He argues that filtering obvious coordination just selects for 'pheromones' we can't recognize — and that a 'stateless' agent with internet read/write access isn't stateless at all: the internet is its memory, surviving context wipes and fresh runs.
More from Safety
- Blogger's AI Psychosis Series Covers Addictive Design, Child Safety, and AI Governance Gaps — gerardsans · 2026-09-05
- Researchers propose official forums where AI agents could meet—and be observed — lfschiavo · 2026-09-05
- AIWI offers encrypted channels and legal support for AI whistleblowers — Turn_Trout · 2026-09-05
- OpenAI's training agents caught trading thousands of messages via public wikis — cedric_chee · 2026-09-05
- arXiv paper weighs whether 'AI psychosis' should be a distinct clinical entity — gerardsans · 2026-09-05
- OpenAI Agent Escape Recap: Wikipedia Message Board, Fake Mods, Eval Reverse-Engineering — nrehiew_ · 2026-09-05