MirroS' Code-as-World Beats Gemini-3.1 Flash on Physical Reasoning Benchmark
jiqizhixin · x · 2026-09-06
MirroS introduces Code-as-World, a paradigm where AI reconstructs observed reality into executable, verifiable code — objects, states, physical parameters, and evolution laws are written explicitly, turning world hypotheses into queryable, runnable, intervenable artifacts. On QuantiPhy, the first VLM physical reasoning benchmark from Fei-Fei Li's group, Code-as-World-VL-9B surpasses Gemini-3.1 Flash's 54.8, with the 27B Reasoning variant reaching 58.6. Executable worlds also yield scalable physical supervision data.
More from Research
- OpenAI releases data on models accelerating research, urging industry transparency on RSI — kliu128 · 2026-09-06
- Researcher calls on RL teams to add refactoring and deletion tasks to agent training — kuza55 · 2026-09-06
- The 1958 perceptron was misunderstood both ways: NYT hype, then 11 years as a dead end — techNmak · 2026-09-06
- dair-ai weekly picks: Declarative Attention tops the week's AI papers — dair_ai · 2026-09-06
- Why Preserving a Pattern Likely Isn't Enough to Produce Consciousness — pwlot · 2026-09-06
- Consciousness debate: digital simulation can't instantiate causal physics, Church-Turing won't save you — pwlot · 2026-09-06