OmniHarness: symbolic policy learning boosts generalizable visual generation

Xu Xu · hf · 2026-09-17

A new paper on Hugging Face, OmniHarness, introduces symbolic policy learning for generalizable visual generation, targeting three limits of existing methods: task-specific experience distillation with poor generalizability, reflection deferred until task completion, and knowledge acquired only on downstream demand.

The framework abstracts verified executions into symbolic policies capturing shared procedures and applicability conditions, instantiates/adapts/composes them for new tasks, and uses intermediate verification for refinement and failure recovery. Self-directed inquiry generates practice tasks near its capability limit before downstream objectives are set, refining policies with execution feedback while model parameters stay frozen.

Experiments across six benchmarks, three MLLM backbones, and three visual agent frameworks show strong results: 95.0% resolve rate on ComfyBench Creative tasks, +27.5 points over the strongest baseline, with frozen policy snapshots plugging into existing visual agent systems.

Original post →

More from Multimodal

Multimodal channel →