VibeWorlding Benchmark: Open 30B Model Beats GPT-5.5 on 3D World Building
Justgototheeffinmoon · reddit · 2026-08-18
A new paper, "VibeWorlding," introduces VWE-BENCH to evaluate if multimodal agents can build interactive 3D open worlds end-to-end from text prompts. The study finds that frontier models like GPT-5.5 and Qwen3.8-Max struggle, with success rates below 60% due to difficulties in precise 3D editing.
Surprisingly, an open 30B parameter model, VibeWorlder-30B-A3B, achieved the best overall Pass@1. The model was post-trained with sandbox tools and verifiers, suggesting that reinforcement learning with verifiable rewards trumps raw parameter scale for this task.
Related event: Tencent's VibeWorlding: Agents Build 3D Open Worlds End-to-End(2 posts)→
More from Multimodal
- Google DeepMind Launches SL2T: Sign-Language-to-Text Model on Pixel 11 — dl_weekly · 2026-08-18
- AI generates live-action Naruto fight: Granny Chiyo vs. Sasori looks insanely real — eyishazyer · 2026-08-18
- ByteDance releases Bernini-Diffusers-v2 with full pipeline — mmowg · 2026-08-18
- AI achieves impossible camera shot: disassembling lens mid-shot — Aiden_Tech_Ai · 2026-08-18
- MiniMax H3 video gen config: LoRA and settings breakdown — erioca · 2026-08-18
- Help: ComfyUI Workflow for Video Re-rendering with Style Transfer — EnvironmentalLime175 · 2026-08-18