Study: Cross-Version Transfer of Qwen Interpretability Lenses
imstilllearningthis · reddit · 2026-08-16
This study tests whether a Jacobian Lens (interpretability tool) fitted for Qwen2.5-72B can be directly applied to Qwen2.5-72B-Instruct without retraining.
Key Findings:
- Reading: The transferred lens performs well in identifying latent entities, with mid-layer rankings even better than the original model. Surface next-token prediction incurs higher costs in deep networks (2x).
- Steering: Direction vectors extracted from the old model (e.g., 'paradox') can successfully remove specific concepts in the new model without breaking descriptive coherence.
Conclusion: Cross-checkpoint transfer is viable, allowing monitoring pipelines to reuse old lenses without refitting for every model update.
More from Models
- 8 hours testing Pangram: anti-AI-detection tricks all failed; only self-typed drafts pass — PawelHuryn · 2026-08-16
- Uncensored Qwen 27B Model Released as GGUF Quantization — BLUECOW009 · 2026-08-16
- Grok 4.6 beats GPT-5.6 Sol on coding agent efficiency, 35% lower cost — rohanpaul_ai · 2026-08-16
- Uncensored vision-enabled Qwen3.6-27B finetune trends on Hugging Face with GGUF release — HauhauCS · 2026-08-16
- Adjusting prediction effort busts cache; watermarking remains — banteg · 2026-08-16
- Anthropic details how Claude’s new watermarks work: mechanism and resistance to editing — TechCrunch AI · 2026-08-16