See2Act teaches robots where to look and how to act in one denoising loop
heghbalz · x · 2026-07-22
A robotics research thread introduces See2Act, a diffusion-based policy for learning where to look while learning how to act.
- The core idea is to couple perception and action inside a single denoising loop.
- The project argues that robots cannot act well on things they cannot see, so gaze selection becomes part of the policy.
- The post links to the paper, project page, and a longer thread explaining the approach.
Because this is a robot method demo, it belongs primarily to robotics/hardware, with research as a secondary angle.
More from Embodied
- China’s driverless delivery trucks are already reshaping logistics — CurieuxExplorer · 2026-07-22
- CHI 2026 best paper uses EMS and embodied AI to guide physical tasks — MacrinePhD · 2026-07-22
- Humanoid robot fight shows may already out-earn many nine-figure startups — kscottz · 2026-07-22
- AlayaRenderer-Flash lifts a generative world renderer from 0.56 FPS to 31.54 FPS — AlayaLab · 2026-07-22
- Autonomous trucks are cutting logistics costs across China’s highways, ports and mines — TansuYegen · 2026-07-22
- A Real Steel-style robot moment is now being framed as reality in China — CurieuxExplorer · 2026-07-22