Boogu-Image-0.1: Open-Sourcing a Unified Multimodal Model

Boogu · hf · 2026-07-16

The Boogu team released **Boogu-Image-0.1**, an open-source family of unified multimodal understanding and generation models, featuring Base, Turbo, Edit, and Edit-Turbo versions. The project focuses on high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual Chinese-English text rendering. The authors note that instead of relying on a single massive model, they achieve significantly better generation and editing results under constrained compute through improved model comprehension, data quality, training pipelines, and **agentic inference-time scaling**. The paper claims it matches or surpasses other open-source models on multiple standard benchmarks and approaches leading closed-source systems. The authors also shared striking cost details: using only **208.62 million unique images**, the theoretical training cost for the base model was roughly **$400,000**. Code, weights, and the training recipe have been open-sourced under Apache 2.0.

Related event: Boogu-Image Open-Sources Unified Multimodal Model Family(2 posts)→

Original post →

More from Multimodal

Multimodal channel →