SceneActBench tests whether VLM agents can act in full 3D scenes
Yifei Zhao · hf · 2026-07-27
- SceneActBench is a benchmark for visually conditioned action on 3D scenes, targeting VLM agents that must act rather than merely describe what they see.
- Existing 3D benchmarks mostly score text answers or single-object operations, leaving full multi-object scene action under-evaluated.
- The benchmark covers five tasks under one unified agent-environment loop, using PNG images or sampled video frames and, when available, supplied 3D assets.
- Final outputs are evaluated against hidden ground truth with task-specific geometric metrics.
- SceneActBench contains 210 source instances and 520 task cases, and the authors report that across 11 proprietary VLMs, overall scores range from 38.6 to 50.2, with no model performing consistently well across tasks.
More from Apps
- Skywork Video says AI video must move from prompting to full production workflows — Shruti_0810 · 2026-07-27
- Skywork Video adds brand guidelines so logos, fonts, and colors can be reused — Shruti_0810 · 2026-07-27
- A small AI app lets readers chat with a book’s “consciousness” — pickover · 2026-07-27
- ChatGPT/Codex Chrome extension is used to build Gmail filters from a prompt — gabrielchua · 2026-07-27
- Built with Claude Code, an LSAT error-log site now lets students share question threads — Isaiah-Burton · 2026-07-27
- TapNow’s CreativeOS powers a Shenzhen AI horror hackathon and pushes video creation into one workflow — 葬AI · 2026-07-27