H3-World Turns MiniMax-H3 into World Model with Only 8K Samples

linoy_tsaban · x · 2026-09-02

H3-World transforms the MiniMax-H3 pretrained language model into a world model without adding new action modules. By converting keyboard controls into textual instructions and injecting them via the model's text pathway, it achieves language-native and temporally grounded control. The method is highly efficient, using only 8,000 gameplay samples, 10,000 LoRA steps, and 0.199% trainable parameters to control character and camera motion, even on unseen action compositions.

Original post →

More from Research

Research channel →