Lab Trains Qwen3-32B to Introspect: Faithful Self-Report Emerges Late, With a Measurable Neural Footprint

davidbau · x · 2026-10-06

David Bau walks through his lab's experiments (@diatkinson) inducing introspection in Qwen3-32B using Dillon Plunkett's Self-Interpretability protocol (arXiv:2505.17120):

Takeaway: the same model can give accurate introspection or parrot-lie about it — and neurons can tell the difference.

Original post →

More from Models

Models channel →