VibeWorlding: Open Source Models Outperform GPT-5.5 in 3D World Building
机器之心 · wechat · 2026-08-26
HKUST(GZ) and Tencent launched VibeWorlding, a multimodal agent framework for building interactive 3D worlds like coding, open-sourcing datasets, a sandbox, and models.
- Components: Includes VWE-Bench (6828 queries, 2616 assets) and VibeWorlding-Gym training framework, introducing a "dual-constraint validator" (physical feasibility + intent satisfaction).
- Results: The open-source VibeWorlder-30B-A3B (based on Qwen3-VL) achieved 59.3% Pass@1 after RL, outperforming GPT-5.5 (57.3%) and Qwen3.8-Max.
- Findings: RL significantly improved 3D reasoning and retrieval, but "precise spatial manipulation" (e.g., collision avoidance) remains a bottleneck for all models. Common failures include coordinate direction errors and over-editing.
More from Multimodal
- Short film showcases Seedance 2.5's realism and directorial control — Uncanny_Harry · 2026-08-26
- Comparing Video Generation: Minimax H3 vs LTX 2.5 — call-lee-free · 2026-08-26
- Thomson Reuters Releases Thomson-1.0-Small Vision-Language Model — thomsonreuters · 2026-08-26
- GPT Image 2 prompt recipe for 100% likeness candid portrait shots — SimplyAnnisa · 2026-08-26
- AI-generated content is ruining cute animals on the internet — Wired AI · 2026-08-26
- ComfyUI workflow automates Minimax H3 video generation and timing benchmarks — GeroldMeisinger · 2026-08-26