HDR: New Diffusion Framework for Multi-Step Visual Reasoning
Zezhong Qian · hf · 2026-07-18
HDR proposes a hierarchical denoising framework for multi-step visual reasoning, organizing video latents into a tree-like hierarchy to perform coarse-to-fine reasoning during generation.
Main Concept
- Coarse-level denoising retains uncertain hypotheses for global planning.
- Fine-level denoising progressively converges these hypotheses into concrete visual states.
- A sparse hierarchical attention mechanism is introduced to reduce attention overhead along the temporal dimension.
Evaluation Results
- The authors constructed a hierarchical, multi-step video reasoning benchmark covering 6 task categories: mazes, Tower of Hanoi, single-line drawing, sliding puzzles, Sokoban, and water pouring, including out-of-distribution samples.
- Compared to the streaming autoregressive diffusion baseline, HDR improved the success rate from 34.22 to 60.29 and average progress from 76.00 to 89.56.
- Inference latency remains at 0.70 seconds per latent, making it 54.2 times faster than bidirectional diffusion.
- Using only 2% of the training data, HDR retains 82.9% of full-data performance, whereas bidirectional diffusion retains only 52.0%.
Additional Results
- In real-world robot experiments, HDR also demonstrated potential for physical interaction and world modeling.
More from Research
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Turning Noise into Signal: Predicting TCR Binding Using AlphaFold3 Hallucinations — quaidmorris · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- PNAS special issue examines copyright, governance, and AI in the legal system — chrmanning · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22