Causal Analysis Reveals Single-Layer Design, Removing 97% of VLA Cross-Attention Params

mathildepapillo · x · 2026-08-05

A research report from Goodfire's Silico platform details optimizing Vision-Language-Action (VLA) model efficiency through causal analysis. The study found that action experts typically read only a narrow band of middle-to-late VLM layers.

Using activation patching on MolmoBot, layer 24 was found to have the largest impact on instructed behavior. Pruning based on this "single-layer design" removed 97% of cross-attention parameters, reducing action expert FLOPs by 93.5% and overall FLOPs by 40.2%. In 1,000 simulation rollouts, the simplified model's success rate remained virtually identical to the original.

Related event: Causal Analysis Optimizes Robot Models, Slashing Compute and Redundant Parameters(2 posts)→

Original post →

More from Embodied

Embodied channel →