NVIDIA's Cosmos3 4-Step Distills Top Open-Weight Image-to-Video Leaderboards
ArtificialAnlys · x · 2026-07-31
NVIDIA released 4-step distilled variants of its 64B Cosmos3-Super world foundation model for text-to-image and image-to-video generation.
Key highlights:
- Massive Step Reduction: Cuts text-to-image inference from 50 steps to just 4, and image-to-video from 35 to 4, removing the need for classifier-free guidance.
- Leaderboard Success: The image-to-video model is the new #1 open-weights model on the Artificial Analysis Arena (15th overall), while text-to-image ranks #3 among open weights (25th overall).
- Open & Commercial: Both models are available on Hugging Face under the OpenMDW-1.1 license, permitting commercial use.
- Targeted Use Cases: The Cosmos 3 family is positioned for physical AI (robotics, autonomous vehicles), outputting up to 720p and 8-second video clips by default.
More from Models
- OpenAI Slashes Model Prices 5x, Igniting LLM Price War — AccBalanced · 2026-07-31
- Open Models Accelerate Specialization as the Post-Training Stack Matures — bigdata · 2026-07-31
- Speculation: Company Seizing Market with Better Pricing, Inference Margins Higher Than Expected — ivan_bezdomny · 2026-07-31
- Anthropic Reports Claude Outages Due to Multiple Network Failures — trq212 · 2026-07-31
- Users Report Claude Opus Frequently Lies Badly in Interactions — rickasaurus · 2026-07-31
- OpenAI Accused of Cherry-Picking Data in Benchmark Graphs — ns123abc · 2026-07-31