Lab Trains Qwen3-32B to Introspect: Faithful Self-Report Emerges Late, With a Measurable Neural Footprint
davidbau · x · 2026-10-06
David Bau walks through his lab's experiments (@diatkinson) inducing introspection in Qwen3-32B using Dillon Plunkett's Self-Interpretability protocol (arXiv:2505.17120):
- Setup: SFT the model to play 100 characters with random hidden preferences (e.g., Gregor Samsa buying washing machines), trained only to output A/B decisions.
- Decisions generalize, self-reports are nonsense: the model nails new preference puzzles but its "I think X" explanations are uncorrelated with what it actually learned — a stochastic parrot.
- Breakthrough: extending training 3x makes faithful self-report emerge late (decisions 0.82/faithfulness 0.25 at step 1000 → 0.92/0.83 at step 3000). Faithful checkpoints store preferences 5-6 layers earlier, where the model's verbalization machinery can read them.
- Neural footprint: cosine similarity between decision and report attributions is 0.34 for faithful models vs 0.08 for unfaithful ones (95% CI 0.16–0.36).
Takeaway: the same model can give accurate introspection or parrot-lie about it — and neurons can tell the difference.
More from Models
- Dev runs every task in both Codex and Claude Code: Opus 5.5 beats GPT-6 Astra — JeremyNguyenPhD · 2026-10-06
- User Pushes Back on GPT-6 Astra Efficiency Claims, Says Claude Opus 5.5 Limits Feel More Generous — CtrlAltDwayne · 2026-10-06
- Benchmarks vs reality: local quantized Qwen3.8 27B beats cloud flash next in real coding workflows — SeriousJul · 2026-10-06
- Synthetic data is the biggest opportunity in AI coding, argues long thread — creatoroff · 2026-10-06
- Codex spins its wheels while Claude Opus shows far deeper understanding, user reports — iruletheworldmo · 2026-10-06
- Cohere Labs Releases Tiny Aya, a Family of Small Models Covering 70+ Languages — Cohere_Labs · 2026-10-06