Oxford VGG Unveils SynCity 3000, Generating Globally Coherent Scene-Scale 3D Worlds
rsasaki0109 · x · 2026-10-02
Researchers at Oxford's Visual Geometry Group (Paul Engstler, Iro Laina, Christian Rupprecht, Andrea Vedaldi) released SynCity 3000, a follow-up to SynCity for text-to-3D scene generation.
- Approach: They adapt an image-to-3D generator into a convolutional operator, fine-tuning it on data from a new synthetic data engine, then apply it convolatively to a dimetric render of the whole scene — producing 3D scenes of arbitrary size and complexity.
- vs. SynCity: tile-based generation leaves visible grid artifacts; SynCity 3000 applies the generator convolatively across overlapping windows at every diffusion step, yielding organic, globally coherent layouts plus fine-grained object placement without rigid grids.
- Results: 100% preference for its layout control over SynCity, 63% for scene quality over SynCity, and 59.3% over the concurrent 3DTown.
- Paper (arXiv), PDF, GitHub, and a World Designer demo are available.
More from Multimodal
- PEARL Debuts First User-History Personalized Image Generation Benchmark, +15% Over Baselines — Bo Ni · 2026-10-02
- First systematic survey of joint video-audio generation and editing: 9 categories, 28 edit types — Abhinav Sharma · 2026-10-02
- Full AI music video for NCT 127 built with Codex, Higgsfield MCP and ComfyUI — leesysysysy · 2026-10-02
- Image editing demo with Nano Banana — tkasasagi · 2026-10-02
- Creator: Opus 5.5 edits videos well, but forcing AI to clip without real need yields garbage — AlchainHust · 2026-10-02
- Reddit user's VEC concept car AI video shows startlingly realistic motion — Vashukanni · 2026-10-02