Boogu-Image-0.1 says 2.08 billion images and $400,000 were enough to reach open-source SOTA

机器之心 · wechat · 2026-07-28

Boogu-Image-0.1, a new open multimodal model family from Huawei Hong Kong Leibniz Lab and university partners, claims open-source first-place results while using only 208 million unique images and about $400,000 in training cost.

The team argues image generation is moving from text-to-image to requirement-to-image, and built the system around stronger text encoding, agentic prompt rewriting, and model routing. The report also shares practical training findings: public benchmarks are increasingly unreliable, structured data can outperform raw scale, some noisy images should be kept if captions describe the defects, and aggressive RL can collapse aesthetic diversity. The whole model, code, and recipe are released under Apache 2.0, with native 2K support and Ascend NPU compatibility.

Original post →

More from Multimodal

Multimodal channel →