Meta's MuseSpark 1.2 Gains Embodied Skills: Robot Navigation, Table Cleaning, and Visual Coding

量子位 · wechat · 2026-08-21

Meta has showcased the multimodal and embodied capabilities of MuseSpark 1.2, demonstrating its ability to control robots and generate interfaces from visuals.

Architecture and Demos:

The system features a two-layer architecture. An upper-layer MuseSpark variant handles high-level planning and task decomposition, while a lower-layer VLA (Vision-Language-Action) model executes physical actions. Demos include a robot searching a room for a duck and a dual-arm robot organizing a desk. The model iteratively refines actions based on real-time visual feedback.

Key Capabilities:

Availability: Currently used internally for media generation and data annotation. Accessible via Meta Model API and MuseCode, with an open-source release confirmed.

Related event: Meta Launches Muse Spark 1.2 with Visual-to-Code and Robot Orchestration Upgrades(9 posts)→

Original post →

More from Embodied

Embodied channel →