30B MoE Embodied Model Activates 3B Parameters

HeyToha · x · 2026-07-12

The post added that this model's training data differs from most video models: in addition to internet video, it incorporates over 70,000 hours of embodied data, including:

The author also noted the use of a 30B MoE architecture, where only 3B parameters are activated during generation, retaining larger capacity at a lower inference cost.

Related event: LingBot-VA/VLA 2.0 Released: Native Embodied Foundation Model(24 posts)→

Original post →

More from Embodied

Embodied channel →