Cross-model training matches self-training: Qwen learns from Llama investigations without privileged access
a_karvonen · x · 2026-09-05
Answering whether privileged access is needed, the authors trained Qwen on investigations of Llama behaviors and vice versa; the cross-trained model matches the self-trained one. However, open-ended self-explanations remain less reliable: consistent generalization appears in only one setting (Qwen-3.5-397B-A17B explaining whether it was influenced by a hint), which the authors flag as an open problem.
Related event: Self-Explanation Training Generalizes Across Models and Tasks(2 posts)→
More from Research
- Delaunay Canopy (ECCV 2026): SOTA building wireframe reconstruction from sparse LiDAR — RexDouglass · 2026-09-05
- Researchers open Postdoc/PhD role on privacy and contextual integrity in AI agents — niloofar_mire · 2026-09-05
- One bit is enough: hidden data encoded via AC polarity flips across systems — MoonL88537 · 2026-09-05
- RNASSTR: new Rfam-based RNA secondary structure dataset with structure-aware train/test splits — chaitjo · 2026-09-05
- Self-explanation training generalizes beyond narrow hint formats to held-out evals — a_karvonen · 2026-09-05
- Two training targets from behavior investigations: counterfactual predictions and open-ended self-explanations — a_karvonen · 2026-09-05