SAE Feature Interpretability Improves Across Layers
Sauers_ · x · 2026-07-14
This repost discusses how jlens is used to describe SAE features, essentially verbalizing highly correlated features like "motor neurons."
By comparing the consistency between jlens-generated feature descriptions and Neuronpedia labels, the author observed that this "verbalization" capability improves with depth. In Qwen3-4b, a significant shift occurs around layer 22. The original post also notes a similar trend in the raw Jacobian results, where later layers exhibit higher Jacobian gains.
More from Research
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22
- NVIDIA says physical AI starts in simulation with OpenUSD and synthetic data — MonaJalal_ · 2026-07-22
- DepthART pushes monocular depth to tiny models at 1000 FPS on RTX A6000 — kwangmoo_yi · 2026-07-22
- Meta says SAM 3 and DINOv3 cut 3D volume labeling from a month to 15 minutes — AIatMeta · 2026-07-22