VibeWorlding Benchmark: Open 30B Model Beats GPT-5.5 on 3D World Building

Justgototheeffinmoon · reddit · 2026-08-18

A new paper, "VibeWorlding," introduces VWE-BENCH to evaluate if multimodal agents can build interactive 3D open worlds end-to-end from text prompts. The study finds that frontier models like GPT-5.5 and Qwen3.8-Max struggle, with success rates below 60% due to difficulties in precise 3D editing.

Surprisingly, an open 30B parameter model, VibeWorlder-30B-A3B, achieved the best overall Pass@1. The model was post-trained with sandbox tools and verifiers, suggesting that reinforcement learning with verifiable rewards trumps raw parameter scale for this task.

Related event: Tencent's VibeWorlding: Agents Build 3D Open Worlds End-to-End(2 posts)→

Original post →

More from Multimodal

Multimodal channel →