4DCodeBench: GPT-6 Astra Max Tops 18 Agents, but More Tokens Don't Buy Better Elo
kwangmoo_yi · x · 2026-10-06
Teams from Stanford, MIT and Johns Hopkins (incl. Joshua Tenenbaum, Alan Yuille, Jiajun Wu) launch 4DCodeBench: coding agents watch a 35s video and write graphics code from scratch to reconstruct 3D geometry, motion and rendering.
Key findings
- GPT-6 Astra [Max] ranks #1 overall, Claude Opus 5.5 [High] close behind; open-weight models trail proprietary ones
- 200 tasks (100 real + 100 simulated videos), 18 agents, scored across five metric families
- Pairwise VLM-as-judge preferences aggregated into Elo, plotted against per-task tokens and cost for a Pareto frontier
- Counterintuitive: more tokens don't consistently yield higher Elo across models; within GPT-6 Astra, higher reasoning effort does improve reconstruction quality
Related event: 4DCodeBench Tests Coding Agents on Rebuilding Dynamic 3D Scenes from Video(4 posts)→
More from coding & agent
- Debate: Coding Agents Can't Compete When Core Harness Abilities Are Locked — pvncher · 2026-10-06
- W&B shows how to turn a production agent failure trace into an eval — wandb · 2026-10-06
- Dev builds Claude Code skill that writes better HTML plans with plain language, mockups and linting — trq212 · 2026-10-06
- OpenAI ships compaction in Responses API, sparking vendor lock-in debate among developers — pvncher · 2026-10-06
- DoorDash launches MCP and CLI for agentic ordering; dev auto-restocks office pantry with camera — Scobleizer · 2026-10-06
- Anthropic ships Claude Code mods: TypeScript functions that rewrite prompts and UI — thione · 2026-10-06