SurgMotion: A New Model for Learning from Surgical Videos

jiqizhixin · x · 2026-07-16

Researchers introduced SurgMotion, a video-native foundation model designed for surgical videos. Instead of relying on pixel-level reconstruction, it directly learns latent motion prediction.

The core idea is to train the model on "motion" rather than "visual details," allowing it to ignore interferences like smoke and reflections to model spatiotemporal action information. The authors highlight three key designs:

Experimentally, SurgMotion achieved new SOTA results across multiple tasks:

Additionally, it covers tasks like skill assessment, polyp segmentation, and depth estimation. The post includes links to the paper, GitHub, Hugging Face, and an explainer report.

Original post →

More from Research

Research channel →