D-OPCD distills agent harness gains into diffusion weights, lifting T2I scores 60.52 to 65.09
Wenxuan Wang · hf · 2026-10-08
Agentic harnesses boost Text-to-Image performance via memory, skills, orchestration, verification and iterative prompt revision — but those gains live outside the diffusion model. D-OPCD (Diffusion On-Policy Context Distillation) treats agent-improved prompts as privileged context and distills the harness's knowledge into the model weights.
Highlights:
- With the proposed Auto Skill Evolver (ASE) T2I agent, D-OPCD raises the average direct-generation score from 60.52 to 65.09 across four benchmarks.
- Once the weights absorb the knowledge, the harness can shed saturated skills and keep evolving: a second ASE round on the updated generator adds another 1.83 points over a skill-free harness, pointing toward co-evolving harness-model T2I systems.
More from Multimodal
- Midjourney prompt shares Kirlian-style glowing aura portraits with --v 8.2 — tisch_eins · 2026-10-08
- VELA 1.0 speeds up MiniMax H3 video gen: 0.8MP render drops from 2:53 to 2:13 — NoMouse9610 · 2026-10-08
- InSpatio-World 1.5 open-sources real-time 4D world model with multi-input support — liuziwei7 · 2026-10-08
- Opus 5.5 drives Blender+Suno via MCP, delivers sync'd 20s masterpiece in 3.5 hours with zero keyframes — sidahuj · 2026-10-08
- User shares strikingly realistic GPT-6 fantasy hunter character portrait — TheUltraBased · 2026-10-08
- LLM-coded Blender beats video models at controllable video generation — jyangballin · 2026-10-08