Meta Launches Muse Image and Muse Video Models

Meta Superintelligence Labs (MSL) officially released its first batch of media generation models: the image model Muse Image and the video model Muse Video. This marks a significant step for Meta in advancing self-developed multimodal capabilities following its AI architecture reorganization.

Core Capabilities and Technical Details

Muse Image is positioned as an agentic model. It can collaborate with Muse Spark to reason about prompts, search the web, and plan before generation. Driven by reinforcement learning, the model utilizes test-time compute to achieve predictable log-linear Elo gains. Furthermore, it supports complex composition from multiple reference images and can write and execute code to accurately generate charts and QR codes.

Evaluation and User Feedback

On the LMArena leaderboard, Muse Image ranks second in both multi-image generation and single-image editing, leading the third-place Nano Banana 2 by 23 points in multi-image generation. In practical tests, user @mark_k noted that its text and detail rendering is excellent and near perfect. However, @spobin conducted a head-to-head comparison with OpenAI's gpt-image-2 and Google's Nano Banana 2 using a 27-point system.

Web Search and Applications

To improve factual accuracy, Muse Image incorporates web search to ground its generations in real-time information. The model is integrated into the Meta AI app and will power Instagram's photo editing features and advertising tools. The previewed Muse Video model also demonstrates strong competitiveness in prompt adherence, visual fidelity, and temporal consistency.

2026-07-08 ~ 2026-07-10 · 71 related posts

4 near-duplicate retellings: Polymarket · rohanpaul_ai · alexandr_wang · rohanpaul_ai