MiniMax-H3 Scores Just 41.97% on a New Omni-Modal Physical World Reasoning Benchmark
Haoyu Zhao · hf · 2026-09-18
A new evaluation framework tests whether the omni-modal generative model MiniMax-H3 can actually reason about the physical world.
- The benchmark is built around four dimensions of physical-world reasoning and exploits omni-modal inputs: implicit prompts with multiple frames, audio-image, prefix-videos, and audio-video combinations—each modality offers only partial evidence, forcing joint cross-modal inference of latent event states and future dynamics.
- Across 517 instances, MiniMax-H3 achieves a 41.97% overall success rate. Video-based Decision Reasoning is strongest at 56.00%, while Audio-based Disambiguation Reasoning reaches only 27.40%.
- The takeaway: effective multimodal integration remains the key bottleneck for unlocking omni-modal inputs. The project is open-sourced at github.com/gulucaptain/MiniMax-H3-Reason.
More from Multimodal
- Recreating videos in MiniMax H3: a 4-step frame-extraction workflow — Negative-Whereas3307 · 2026-09-18
- AI artist 'Will' drops a new song daily, with persistent memory and personality — kun101 · 2026-09-18
- Kijai's new MH3 VAE cuts VRAM use: 1MP 10s video now runs on an RTX 3060 12GB — Chiduk99 · 2026-09-18
- MiniMax Design adds custom model support, 20% off annual plans — Hailuo_AI · 2026-09-18
- Demo: GPT-6 Astra edits video in DaVinci via MCP at impressive speed — toolstelegraph · 2026-09-18
- Structured JSON prompt recreates Tomie character with Nano Banana Pro — creatoroff · 2026-09-18