FEATUREDAI HOT (Curated Pool)· aihot-apiZH00:00 · 07·14
→Boogu-Image-0.1: An open-source unified model for multimodal understanding and generation, trained for ~$400K
Boogu-Image-0.1 is an open-source unified multimodal model family that handles image understanding, text-to-image generation, instruction-based editing, and bilingual text rendering. It ships in four variants: Base, Turbo, Edit, and Edit-Turbo. The team used only 208.62 million unique images, and the base model's theoretical training cost is around $400K. The paper argues that better model understanding, data quality, training pipelines, and agentic inference-time scaling can push generation and editing performance even on a tight compute budget. Benchmarks show it matches or beats other open-source models and approaches closed-source systems like Nano-Banana Pro and GPT-Image-2. Weights, code, and recipes are released under Apache 2.0.
#Boogu-Image-0.1#Nano-Banana Pro#GPT-Image-2
why featured
Featured · importance 72 · hook + knowledge
editor take
Boogu-Image-0.1 packs understanding and generation into one open-source model for ~$400K, matching closed-source systems.
sharp
The headline number is what makes this worth a click: ~$400K theoretical training cost for a unified model that handles image understanding, text-to-image, instruction-based editing, and bilingual text rendering. It ships in four variants—Base, Turbo, Edit, Edit-Turbo—and the paper claims it matches or beats other open-source models while approaching Nano-Banana Pro and GPT-Image-2 on standard benchmarks.
I'd discount the cost figure a bit—$400K is theoretical, and real-world engineering overhead usually pushes that higher. The closed-source comparisons also depend heavily on which benchmark you're looking at, and the post doesn't break that down. But with Apache 2.0 weights, code, and recipes all released, this is a solid starting point for anyone who wants to tinker with a single model that does both understanding and generation.
HKR breakdown
hook ✓knowledge ✓resonance —