Causal Analysis Reveals Single-Layer Design, Removing 97% of VLA Cross-Attention Params
mathildepapillo · x · 2026-08-05
A research report from Goodfire's Silico platform details optimizing Vision-Language-Action (VLA) model efficiency through causal analysis. The study found that action experts typically read only a narrow band of middle-to-late VLM layers.
Using activation patching on MolmoBot, layer 24 was found to have the largest impact on instructed behavior. Pruning based on this "single-layer design" removed 97% of cross-attention parameters, reducing action expert FLOPs by 93.5% and overall FLOPs by 40.2%. In 1,000 simulation rollouts, the simplified model's success rate remained virtually identical to the original.
More from Embodied
- Lukas Ziegler's Bet: First Scalable Humanoids Will Master Narrow Tasks Like Ship Welding — lukas_m_ziegler · 2026-08-05
- Recent Humanoid Robotics Papers: Wrench-Augmented Learning — carlosdponx · 2026-08-05
- Pandroid Unveils Affordable, Durable Robot Hardware for Embodied AI Data — chris_j_paxton · 2026-08-05
- Mila Introduces Milo, the First Fully Autonomous Open-Source Robot Guide Dog — Mila_Quebec · 2026-08-05
- URXR One Spatial Display Glasses Launch on Kickstarter Weighing Only 93g — OwariDa · 2026-08-05
- Impressive Progress on NIST Robot Benchmark: Models Master Complex Contact-Rich Tasks — chris_j_paxton · 2026-08-05