Boogu-Image-0.1: Open-Sourcing a Unified Multimodal Model

Boogu · hf · 2026-07-16

The Boogu team released Boogu-Image-0.1, an open-source family of unified multimodal understanding and generation models, featuring Base, Turbo, Edit, and Edit-Turbo versions.

The project focuses on high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual Chinese-English text rendering. The authors note that instead of relying on a single massive model, they achieve significantly better generation and editing results under constrained compute through improved model comprehension, data quality, training pipelines, and agentic inference-time scaling. The paper claims it matches or surpasses other open-source models on multiple standard benchmarks and approaches leading closed-source systems.

The authors also shared striking cost details: using only 208.62 million unique images, the theoretical training cost for the base model was roughly $400,000. Code, weights, and the training recipe have been open-sourced under Apache 2.0.

Related event: Boogu-Image Open-Sources Unified Multimodal Model Family(2 posts)→

Original post →

More from Multimodal

Multimodal channel →