Tencent Hunyuan Launches Hunyuan3D-Buffalo 1.0 for Unified 3D Generation, Understanding, and Editing

Tencent-Hunyuan · hf · 2026-08-05

Tencent Hunyuan has introduced Hunyuan3D-Buffalo 1.0, a unified framework that integrates 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture.

To overcome the scarcity of 3D multimodal data, the team constructed an 87M-scale corpus comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs. Architecturally, it combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis.

Experiments demonstrate that the model achieves state-of-the-art performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Analysis further confirms that unifying generation and understanding effectively improves editing performance.

Original post →

More from Multimodal

Multimodal channel →