4DCodeBench: new benchmark tests coding agents on reconstructing dynamic 3D scenes from video
CSProfKGD · x · 2026-10-06
Researchers from Stanford, JHU and elsewhere (authors include Jiajun Wu, Joshua Tenenbaum, Alan Yuille) introduce 4DCodeBench, asking whether coding agents can watch a video of a physical event and write, from scratch, graphics code that reconstructs the scene's 3D geometry, its motion over time, and the rendering back into video.
- Scale: 200 tasks (100 real + 100 simulated videos), ranking 18 agents.
- Evaluation: appearance, geometry and motion are scored across five metric families; VLM-as-judge pairwise preferences are aggregated into Elo ratings.
- Leaderboard: GPT-6 Astra [Max] leads, with Claude Opus 5.5 [High] close behind; open-weight models generally trail proprietary ones.
- Compute findings: across models, higher token use does not consistently yield higher Elo; within GPT-6 Astra, raising reasoning effort improves reconstruction quality at the cost of more tokens.
Related event: 4DCodeBench Tests Coding Agents on Rebuilding Dynamic 3D Scenes from Video(4 posts)→
More from coding & agent
- TanStack Launches MCP Server for AI-Assisted Docs Search and Scaffolding — modelcontextprotocol · 2026-10-06
- SnapRender Ships MCP Server to Capture Website Screenshots via AI Agents — modelcontextprotocol · 2026-10-06
- SFT then RL doesn't fix agent looping: 29% of runs hit turn cap vs 0% for RL alone — VikParuchuri · 2026-10-06
- RL Post-Training Eliminates Agent Tool-Call Loops: 92% Loop Rate Drops to 0 — VikParuchuri · 2026-10-06
- RL Alone Eliminates Agent Loops: 0% at Temp 0 vs 92% for SFT, Datalab Finds — VikParuchuri · 2026-10-06
- Coding agents cause runaway scope creep in research, warns scientist — rishabh16_ · 2026-10-06