Boogu-Image-0.1: Open-Sourcing a Unified Multimodal Model
Boogu · hf · 2026-07-16
The Boogu team released **Boogu-Image-0.1**, an open-source family of unified multimodal understanding and generation models, featuring Base, Turbo, Edit, and Edit-Turbo versions. The project focuses on high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual Chinese-English text rendering. The authors note that instead of relying on a single massive model, they achieve significantly better generation and editing results under constrained compute through improved model comprehension, data quality, training pipelines, and **agentic inference-time scaling**. The paper claims it matches or surpasses other open-source models on multiple standard benchmarks and approaches leading closed-source systems. The authors also shared striking cost details: using only **208.62 million unique images**, the theoretical training cost for the base model was roughly **$400,000**. Code, weights, and the training recipe have been open-sourced under Apache 2.0.
Related event: Boogu-Image Open-Sources Unified Multimodal Model Family(2 posts)→
More from Multimodal
- A prompt for Magnific and GPT-2 produced a dense surreal comic about unlived lives — CurieuxExplorer · 2026-07-21
- Seedance 2.0 prompt aims for MiniDV-style footage with handheld imperfections — eyishazyer · 2026-07-21
- AI-made “cat mode” stunt turns a skateboard clip into a surreal landing demo — taherdhanera · 2026-07-21
- OCT-Bench sets 10,076 questions to test whether multimodal models really understand retinal scans — Baochen Fu · 2026-07-21
- LTX-2.3 face-and-voice LoRA training can work on 12GB VRAM with heavy tradeoffs — __alpha_____ · 2026-07-21
- Seedance 2.0 turns one reference image into a cinematic fight scene — techhalla · 2026-07-21