Boogu-Image-0.1 says 2.08 billion images and $400,000 were enough to reach open-source SOTA
机器之心 · wechat · 2026-07-28
Boogu-Image-0.1, a new open multimodal model family from Huawei Hong Kong Leibniz Lab and university partners, claims open-source first-place results while using only 208 million unique images and about $400,000 in training cost.
The team argues image generation is moving from text-to-image to requirement-to-image, and built the system around stronger text encoding, agentic prompt rewriting, and model routing. The report also shares practical training findings: public benchmarks are increasingly unreliable, structured data can outperform raw scale, some noisy images should be kept if captions describe the defects, and aggressive RL can collapse aesthetic diversity. The whole model, code, and recipe are released under Apache 2.0, with native 2K support and Ascend NPU compatibility.
More from Multimodal
- Netflix’s ID-V2V preserves identity while restyling videos from one source clip — netflix · 2026-07-28
- dRAE scales visual tokenization to 131,072 codes without codebook collapse — burny_tech · 2026-07-28
- New paper maps compute-optimal scaling laws for native multimodal pre-training — burny_tech · 2026-07-28
- New ComfyUI node converts audio into MIDI for music workflows — MuziqueComfyUI · 2026-07-28
- Exploring ComfyUI Basics: Why Separate Checkpoint and KSampler in Workflows? — DavidThi303 · 2026-07-28
- Running Z Image Turbo on RTX 4060: How to Break Through the Quality Ceiling? — Dangerous_Ring_435 · 2026-07-28