VLX-VR video reasoning model tops MINERVA at 78.8%, passing a real mechanical physics check
卡尔的AI沃茨 · wechat · 2026-09-10
The author stress-tested VLX-VR, a new video deep-reasoning model from OmAI (联汇), using a GPT6-generated 3D animation of Cessna 337 landing-gear retraction.
- Mechanical verification: the model accurately described the full drivetrain, gear-rack meshing, hydraulic actuator logic and over-center self-locking, concluding the demo had no physics errors — without inventing flaws to please the prompt.
- Causal reasoning: asked what would happen if the self-lock failed at minute 2, it reasoned forward along the timeline, showing it tracks cause-and-effect between actions.
- Benchmarks: scores 78.8% on MINERVA (1,000+ long-video questions with auditable reasoning traces), ranking first publicly — vs o1 at 43.48%, GPT-4o at 45.54%, and Gemini 2.5 Pro at 57.97% on 15+ minute videos. Step-consistency on correct answers is 96.2%, and accuracy rises with video length (76.70%/78.73%/80.92%). Human baseline: 92.54%.
- Weaknesses: counting dense tiny parts and catching faint local deformations.
- More tests: on a 9x-speed screen recording of industrial CAD modeling it caught an AttributeError at second 11, identified PythonJournal script automation, and counted 22 solids with exact dimensions; on a four-way robot-arm experiment video it distinguished human teleoperation from robot execution (noting 1s pauses).
- Product: the OttoBox video agent built on VLX-VR locates shots in massive footage libraries and produces rough cuts in 30 minutes (vs 8-10 hours manually), running fully on-device.
More from Multimodal
- Creator makes 2D electro-pop anime music video with just a prompt using MiniMax H3 — Hailuo_AI · 2026-09-11
- Street View to driving footage: GPT Astra fetches images, MiniMax H3 turns them into dashcam video — Hailuo_AI · 2026-09-11
- Single-author ECCV 2026 paper makes rolling shutter correction practical — ducha_aiki · 2026-09-11
- AI digital human covers Japanese classic so realistically viewers can't tell — JourneymanChina · 2026-09-11
- ComfyUI Style Explorer Adds LoRA Preview Catalog and Sharing — neonsparksuk · 2026-09-11
- 4 favorite Midjourney V6.1 --sref style codes, ready to copy — michaelrabone · 2026-09-11