OmniHarness: symbolic policy learning boosts generalizable visual generation
Xu Xu · hf · 2026-09-17
A new paper on Hugging Face, OmniHarness, introduces symbolic policy learning for generalizable visual generation, targeting three limits of existing methods: task-specific experience distillation with poor generalizability, reflection deferred until task completion, and knowledge acquired only on downstream demand.
The framework abstracts verified executions into symbolic policies capturing shared procedures and applicability conditions, instantiates/adapts/composes them for new tasks, and uses intermediate verification for refinement and failure recovery. Self-directed inquiry generates practice tasks near its capability limit before downstream objectives are set, refining policies with execution feedback while model parameters stay frozen.
Experiments across six benchmarks, three MLLM backbones, and three visual agent frameworks show strong results: 95.0% resolve rate on ComfyBench Creative tasks, +27.5 points over the strongest baseline, with frozen policy snapshots plugging into existing visual agent systems.
More from Multimodal
- Videoclaw launches as a Mac video agent: prompt ChatGPT or Claude to edit and generate video — alifcoder · 2026-09-17
- StepFun and ACE Studio launch music foundation model StepAudio 3 Music — StepFun_ai · 2026-09-17
- StepFun and ACE Studio launch StepAudio 3 Music, a song-generation foundation model — StepFun_ai · 2026-09-17
- Fan recreates Blue Archive Shiroko scene with AI animation, tutorial coming — Wakaiko · 2026-09-17
- UK dev quit over coding agents, built Videoclaw: chart image to narrated video for $0.39 — JaynitMakwana · 2026-09-17
- Wan2GP trick: inject character reference sheets as the first frame for consistent generations — protocol-apps · 2026-09-17