NVIDIA, MIT and Oxford's Physis-Lang tops Physics-IQ with 48.2, beats Veo 3.1 on physics
机器之心 · wechat · 2026-10-09
NVIDIA, MIT and Oxford introduced Physis-Lang, a physics-aware language representation that injects physical causes, laws, and outcomes into video captions across data curation, training, and inference—forcing video models to fill in the intermediate physics of events like melting butter.
Key components
- Physics-enriched captions at training and inference, plus negative descriptions to suppress implausible behavior;
- PhysCapBench: 246 videos with 3,794 human-verified atomic assertions, scored via Precision/Recall/F1;
- Self-evolving shared caption prompts lifted PhysCapBench F1 from 78.64 to 87.82; fixing the model and only changing inference captions raised PhyGenBench from 64.17 to 67.29;
- Model weaknesses are turned into retrieval labels to mine real training videos, yielding a 183K-video dataset with 8-point gains on thermal/chemical categories of VideoPhy-2.
Results: Cosmos3-Super and Nano scored 48.2 and 43.3 on Physics-IQVerified (top two), Super beating the previous best by 5.5 points; open-source SOTA on four physics benchmarks, surpassing Veo 3.1 on three.
More from Multimodal
- One prompt turned OpenAI's 722 math papers into a 92-second animated explainer via Opus — FinanceYF5 · 2026-10-10
- Dev recreates fictional ChatGPT UIs for Windows 98, XP, Vista and 7 — Midnight_Sun_BR · 2026-10-10
- MiniMax H3 removes people from videos inside ComfyUI — RobbaW · 2026-10-10
- BudgetPix: Google and UIUC's pixel diffusion model adapts compute per image, cutting tokens to 10% — CSProfKGD · 2026-10-10
- Free Female Face Prompt Builder adds natural-language-to-SD prompt conversion — dobelmont · 2026-10-10
- Model tuning could stop AI video from fighting 2D animation style, says Andrew Carr — andrew_n_carr · 2026-10-10