H3-World turns MiniMax H3 into a controllable world simulator with 8k gameplay samples
sachasayan · reddit · 2026-09-02
H3-World repurposes MiniMax H3's pretrained text pathway into language-native world control: character and camera actions are composed as textual instructions and injected per video latent interval for temporally grounded control.
Remarkably efficient — just 8,000 gameplay samples, 10,000 LoRA steps, and 0.199% trainable parameters yield controllable character/camera motion that generalizes to unseen action compositions and visual scenarios. Paper (arXiv 2609.01560), code, model and project page are all public.
More from Multimodal
- Controlled Z-Image Base experiments show explicit composition prompts reshape the whole frame — Maleficent-Bowl-4841 · 2026-09-02
- NVIDIA Unveils DLSS 5: 3D-Guided Neural Rendering Called Biggest Visual Leap Since 3D Itself — ctnzr · 2026-09-02
- Sprite Fusion generates production-ready AI pixel art at native sizes in seconds — HugoDuprez · 2026-09-02
- Rewriting an infographic prompt as a data spec: 12 to 20 characters on SenseNova U1.5 Lite — Pristine_Weight_4705 · 2026-09-02
- H3 Motion Context 0.5.0 chains MiniMax H3 clips with motion and soundtrack carryover — Sad_Berry_4621 · 2026-09-02
- AI-generated cat cartoon 'Whiskerhold' nails the Sunday-morning animation vibe — MosskeepForest · 2026-09-02