HiDream launches native omni-modal video model, debuts top-10 on both leaderboards
新智元 · wechat · 2026-09-22
HiDream.ai released HiDream-O1-Video-1.0, a native omni-modal video model that debuted at #4 on Artificial Analysis' image-to-video leaderboard (with audio) and #8 on the Arena blind test, joining MiniMax, ByteDance Seedance and Alibaba Wan in the top tier.
- Focus on physical plausibility: demos cover Doppler effects, free fall, water balls in space, and material-specific collision behavior; native 5–20s generation with duration driven by narrative, not fixed prompts
- A multimodal intent-understanding module converts colloquial prompts into structured shot planning (scene, characters, lighting, sound design)
- Joint text-video-audio modeling for native A/V sync, including Chinese lip sync and emotional shifts
- Post-training uses Diffusion RL with multimodal rewards, backed by three papers: DMSampler (ICML 2026, cheaper sampling), DiffusionCompass (ECCV 2026, reward-hacking resistance), EvoID (CVPR 2026, identity preservation)
- Part of HiDream's UiT foundation spanning image, interactive world model, embodied model and video. Note: promotional framing — physical capability claims rest on official demos
More from Multimodal
- Refining MiniMax H3 720p video to 2K on an 8GB GPU: a ComfyUI workflow discussion — ashendonep · 2026-09-23
- Tencent Hunyuan's Hy Image 3.5 preview lands in ComfyUI with +30% human eval win rate — TencentHunyuan · 2026-09-23
- Anyone can now generate videos like these — full prompt shared — umesh_ai · 2026-09-23
- FlovaAI launches year-long price lock for Seedance 2.5 at $0.044/sec in 720p — thetripathi58 · 2026-09-23
- Qwen-Image-2.1 4-step turbo LoRA generation demo — CodeMichaelD · 2026-09-23
- "No amount of graphics made a bad game good": devs on 3D generation in game design — rms80 · 2026-09-23