HiDream's interactive world model tops WBench benchmark
机器之心 · wechat · 2026-08-17
HiDream.ai has released HiDream-O1-World, the world's first natively full-modal interactive world model. Built on the proprietary UiT architecture, it supports text, image, and interaction inputs to generate dynamic worlds with long-term spatio-temporal and physical consistency, allowing for real-time roaming and editing.
Core Capabilities
- Roaming Mode: Supports free switching between first/third-person views with real-time, consistent scene updates.
- Editing Mode: Allows real-time control of character actions or environmental changes (e.g., weather) based on physical simulation.
- Multi-Style Support: Covers photorealistic, anime, and AAA game styles.
Benchmark Performance
In the WBench benchmark by Meituan and Fudan University, HiDream-O1-World topped the Navi leaderboard on its first attempt. It scored 73.3 in Physical consistency (1st) and 88.0 in Consistency (Overall 1st), outperforming models like Tencent Hunyuan 1.5 and Alibaba HappyOyster.
Technical Architecture
- UiT (Unified Transformer): Abandons traditional splicing, mapping image pixels, text tokens, and video voxels into a shared space for unified interaction within a single Transformer network.
- Geometry-then-Appearance: A two-stage framework that infers geometry first to ensure spatial consistency, then generates appearance, resolving distortion during large view changes.
- Memory 3D + TTT: Uses spatial memory injection and Test-Time Training to maintain scene memory and physical logic during long-term interactions.
Applications
Beyond interactive movies and gaming, HiDream is partnering with Noitom Robotics to generate large-scale training data for embodied intelligence, addressing the gap between real-world data costs and simulation fidelity.
More from Multimodal
- Seedance 2.5 Tops Multi-Image-to-Video Benchmark with Elo 1400 — rohanpaul_ai · 2026-08-17
- SPARGen unifies spatial perception and reasoning via multimodal generation — Jinsheng Quan · 2026-08-17
- Marionette predicts world states and renders geometry via diffusion — Zian Meng · 2026-08-17
- 2D moments into 3D holograms: the geospatial memory palace comes to life — bilawalsidhu · 2026-08-17
- Fell back to LTX 2.3 for Tony Soprano parody after 2.5 failed to recognize him — Unluckiestfool · 2026-08-17
- MiniMax H3 Hands-On: 8 Commercial Use Cases, Cost-Effective Anime-Style Video Generation — 卡尔的AI沃茨 · 2026-08-17