Cross-model training matches self-training: Qwen learns from Llama investigations without privileged access

a_karvonen · x · 2026-09-05

Answering whether privileged access is needed, the authors trained Qwen on investigations of Llama behaviors and vice versa; the cross-trained model matches the self-trained one. However, open-ended self-explanations remain less reliable: consistent generalization appears in only one setting (Qwen-3.5-397B-A17B explaining whether it was influenced by a hint), which the authors flag as an open problem.

Related event: Self-Explanation Training Generalizes Across Models and Tasks(2 posts)→

Original post →

More from Research

Research channel →