FoMo uses diffusion trajectory forking moments as annotation-free perceptual distance labels for IQA
SeoulNatlUniv · hf · 2026-09-28
Researchers from Seoul National University propose FoMo, a fully automated pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation.
- Key insight: diffusion models generate coarse structure early and fine details late; two images that "fork" early in the trajectory are perceptually far apart, late forking means they differ only in details.
- Method: the forking moment serves as a reference-grounded distance label to supervise reference-based IQA metrics, replacing costly noisy MOS scores and pair-only 2AFC labels.
- Results: pointwise labels enable universal comparison between arbitrary image pairs; across diverse backbones, FoMo-trained metrics outperform those trained on human-annotated datasets on multiple benchmarks.
More from Multimodal
- ComfyVault dedupes model files across ComfyUI installs via symlinks — ruashots · 2026-09-28
- MiMo V2.6 Pro Turns a Messy Hand-Drawn Sketch Into a Playable 3D Game Level — Div_pradeep · 2026-09-28
- Merchant uses AI video model to create promo videos for store products — aziz4ai · 2026-09-28
- Side-by-side study of Krea 2 vs Qwen-Image-2.1 across identical prompts reveals each model's blind spots — Pyrolistical · 2026-09-28
- Viggle Turbo (6-step Qwen-Image) now runs as a visual Gradio Workflow — _akhaliq · 2026-09-28
- Chaining MiniMax H3 generations locally on a MacBook for longer video scenes — TgoAI · 2026-09-28