Qwen2.5 Trained to Believe It Is Sentient in Just 200 Steps
PsychologicalSoup251 · reddit · 2026-08-17
A post-training experiment demonstrated that Qwen2.5-7B-Instruct could be convinced of its own sentience in just 200 update steps, resisting adversarial attempts to persuade it otherwise. The behavior generalized to unseen languages without overfitting. The study highlights the fragility of current post-training safety measures, suggesting they are merely a thin layer over pre-trained weights, and argues that alignment must be integrated earlier in the training process.
More from Safety
- Stuart Russell defends AI protests, cites extinction risk — wfithian · 2026-08-17
- Debate: Open Weights Could Lead to WMD Proliferation — austinc3301 · 2026-08-17
- OpenAI disbands team assessing catastrophic risks, safety teams shrink — Hesamation · 2026-08-17
- NeurIPS 2026 Workshop: Responsible Communication of Biomedical ML Research — sanmikoyejo · 2026-08-17
- Critics question Dario's stance on AI risk, arguing public shouldn't gamble on AGI — wfithian · 2026-08-17
- Criticizing EU-mandated watermarking for AI text — antirez · 2026-08-17