ToolArtist: Agentic Image Generation via Unified Multimodal Models
Jiahao Zhao · hf · 2026-08-06
To overcome T2I limitations in complex reasoning and external knowledge, this paper introduces ToolArtist, a fully agentic image generation model post-trained from a Unified Multimodal Model (UMM).
- Unified Policy: Dynamically orchestrates reasoning, external tool use (like search), and native image generation within a single policy.
- Training: Uses a teacher agent for SFT trajectory collection and introduces RAD-GRPO during RL to jointly optimize intent and quality rewards.
- Results: Placing the entire open-world image generation process under an agent policy consistently outperforms fixed pipelines. Training data and infrastructure are open-sourced.
More from Multimodal
- Advanced MiniMax H3 Filmmaking Workflow: From Prompts to Editing — Tricky_Algae2625 · 2026-08-06
- Testing LTX upscaler: Generating high-res images on low VRAM GPUs — AniZeee · 2026-08-06
- Elon Musk Announces Launch of Grok Imagine Image Generation — elonmusk · 2026-08-06
- Meta Unpacks Multimodal Pretraining: Strong Generation with 5% Compute — facebook · 2026-08-06
- UniWorld-View: Large-Baseline Novel View Synthesis via Video Diffusion — Haiyang Zhou · 2026-08-06
- HelloWorld: Enabling Real-Time Social Interaction in Video World Models — Liangyang Ouyang · 2026-08-06