AutoRef Auto-Optimizes Image Generation Harnesses, Boosting FLUX.2 to Match Proprietary Models
Yuta Oshima · hf · 2026-09-30
AutoRef: Automatically Optimizing Harnesses for Multi-Reference Image Generation
Multi-reference image generation remains hard: models omit or duplicate subjects, or paste them together unnaturally. Agent-based pipelines (generator + reasoning model + executable harness) help, but hand-written harnesses vary widely in quality.
Method: AutoRef keeps both models frozen and uses a coding agent to iteratively rewrite the harness code. It separates tasks whose feedback drives proposals from tasks used for candidate selection, and continues search from a beam of top-ranked harnesses.
Results:
- The discovered AutoRef-Harness lifts open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 on held-out four-reference tasks of the MultiBanana benchmark, matching or exceeding proprietary models like Nano Banana Pro and GPT-Image-1.5
- Without re-optimization, the same harness transfers across different generators, reference counts, benchmarks, evaluators, and reasoning models
More from Multimodal
- Second Claude Opus Greek-art experiment: score again unprompted, done in ~20 minutes — Kyrannio · 2026-09-30
- Claude Opus crafts Greek art-style piece with original score in 30 minutes — Kyrannio · 2026-09-30
- Scenario generates isometric walk and attack cycles, hailed as most exciting AI game-dev tool — AIandDesign · 2026-09-30
- Open-source YeU2 LoRA delivers raspy 80s-90s female rock-soul vocals — -becausereasons- · 2026-09-30
- CompVis improves Distributional Diffusion Models: 4.48 FID at 4 steps on ImageNet — CompVis · 2026-09-30
- NeurIPS paper to Opus 5.5 music video pipeline explores the "Interviewer Effect" — jankulveit · 2026-09-30