3DHarnessBench probes agentic 3D-to-code skills of frontier VLMs
ftm_guney · x · 2026-09-09
Researchers from Institut Polytechnique de Paris introduce 3DHarnessBench, a benchmark evaluating how frontier vision-language models reconstruct 3D geometry as Blender Python code. Four harness settings progressively grant agent access: single-view, multi-view, active visual (arbitrary viewpoints), and full 3D interaction via Blender MCP function calls.
- All frontier models improve significantly with richer function-call access, but gains are strongly model-dependent, revealing uneven agentic 3D-to-code capabilities
- The benchmark probes visual perception, active inference, tool calling, and self-correction
- Benchmark, code, outputs, and agent trajectories will be released for reproducibility
More from coding & agent
- Demo: install agentgateway on a Raspberry Pi in under 60s to track OpenAI token spend — bibryam · 2026-09-09
- ByteDance, Alibaba and others split between work and personal AI assistants; memory may decide — sujingshen · 2026-09-09
- With: a new open-source systems language compiling to native code via LLVM — QuixiAI · 2026-09-09
- NVIDIA launches CUDA Rust: two tracks to write GPU kernels in Rust — ducha_aiki · 2026-09-09
- Agent Astra builds a Grim Dawn save, streams via Moonlight, and plays the 3D game in real time — IridiumEagle · 2026-09-09
- pstack 0.15.0 cuts token usage 3-11% and adds two new agent principles — RachelVT42 · 2026-09-09