Boogu Releases Open-Source Multimodal Model Rivaling Closed Systems at $400K Cost

jiqizhixin · x · 2026-08-08

Boogu Team has released Boogu-Image-0.1, an open-source multimodal model capable of both understanding and generating images, rendering bilingual text, and performing instruction-based image edits.

The core approach focuses on strengthening image understanding through better encoders, smarter prompt rewriting, and improved training data to dramatically boost generation quality. Despite a minimal compute budget, it matches or beats top open-source models and closely approaches closed systems like Nano-Banana Pro and GPT-Image-2.

Training Details:

Weights, code, and training recipes are now open-sourced under the Apache 2.0 license.

Original post →

More from Multimodal

Multimodal channel →