OmniHarness: visual agents that self-practice and bank verified workflows hit 95% on creative tasks

新智元 · wechat · 2026-09-18

Researchers from Beihang, CUHK and NUS released OmniHarness, a framework where visual agents practice autonomously before downstream tasks arrive, verify results, and distill successful workflows into reusable strategies. It reaches 95% on ComfyBench creative tasks, beating the best baseline SymbOmni by 27.5 points; with GPT-4o it scores 89.5% overall, rising to 92.5% with Codex planning.

The pipeline has three parts: task selection scores 10 candidates per round by novelty × capability boundary, favoring challenging-but-tractable tasks; execution compiles plans into ComfyUI workflows with per-step checks and localized repair; accumulation abstracts verified flows into a strategy library (templates, applicability conditions, resource dependencies) while failures log evidence and fixes, with Wilson-interval reliability updates.

The frozen strategy library transfers across agents without fine-tuning: ComfyAgent jumps 32.5%→57%, ComfyMind 83%→88%, SymbOmni→89%, with the largest gains on creative tasks. Image-trained strategies also adapt to video workflows (85.9% overall). Costs come from autonomous practice and multi-round verification. Paper: arxiv.org/abs/2609.16057.

Related event: OmniHarness lets visual agents self-practice, hitting 95% on creative tasks(2 posts)→

Original post →

More from Multimodal

Multimodal channel →