Boogu-Image-0.1: Open-Sourcing a Unified Multimodal Model
Boogu · hf · 2026-07-16
The Boogu team released Boogu-Image-0.1, an open-source family of unified multimodal understanding and generation models, featuring Base, Turbo, Edit, and Edit-Turbo versions.
The project focuses on high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual Chinese-English text rendering. The authors note that instead of relying on a single massive model, they achieve significantly better generation and editing results under constrained compute through improved model comprehension, data quality, training pipelines, and agentic inference-time scaling. The paper claims it matches or surpasses other open-source models on multiple standard benchmarks and approaches leading closed-source systems.
The authors also shared striking cost details: using only 208.62 million unique images, the theoretical training cost for the base model was roughly $400,000. Code, weights, and the training recipe have been open-sourced under Apache 2.0.
Related event: Boogu-Image Open-Sources Unified Multimodal Model Family(2 posts)→
More from Multimodal
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Non-coder builds full-featured Android ComfyUI client with ChatGPT, submits to Google Play — ComfierUI · 2026-09-11
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11