Boogu Releases Open-Source Multimodal Model Rivaling Closed Systems at $400K Cost
jiqizhixin · x · 2026-08-08
Boogu Team has released Boogu-Image-0.1, an open-source multimodal model capable of both understanding and generating images, rendering bilingual text, and performing instruction-based image edits.
The core approach focuses on strengthening image understanding through better encoders, smarter prompt rewriting, and improved training data to dramatically boost generation quality. Despite a minimal compute budget, it matches or beats top open-source models and closely approaches closed systems like Nano-Banana Pro and GPT-Image-2.
Training Details:
- Trained on only 208 million images.
- Total training cost of roughly $400K.
Weights, code, and training recipes are now open-sourced under the Apache 2.0 license.
More from Multimodal
- Describe your dream world to an AI dragon, which generates the planet for you — repligate · 2026-08-24
- Using kintsugi texture to fix cracks in edited 3D meshes — repligate · 2026-08-24
- Generating Hannibal Character Videos with FL2VA Model — Nimblecloud13 · 2026-08-24
- MiniMax H3 Revives Medieval Short Stories: Complete Workflow Shared — zanatas · 2026-08-24
- NAPE Audio Pretraining Achieves SOTA Without Decoders — kastnerkyle · 2026-08-24
- H3 excels at generating complex space scenes — SIR_NVAX_A_LOT · 2026-08-24