Meta Unveils Muse Spark 1.2: Vision-to-Code, Robot Navigation, Audio-Visual Understanding

AIatMeta · x · 2026-08-21

Meta officially announced Muse Spark 1.2, supporting a broad range of multimodal tasks: turning visuals into working code, translating perception into physical action, and robust audio-visual understanding for video-heavy enterprise workflows.

Alongside new evals, Meta shared demos of the model's visual understanding and reasoning — starting with one where Muse Spark parses multimodal observations and calls tools to guide a robot through an unstructured environment to find a rubber duck.

Related event: Meta Unveils Muse Spark 1.2 with Visual-to-Code and Robot Orchestration(7 posts)→

Original post →

More from Models

Models channel →