SurgMotion: A New Model for Learning from Surgical Videos
jiqizhixin · x · 2026-07-16
Researchers introduced SurgMotion, a video-native foundation model designed for surgical videos. Instead of relying on pixel-level reconstruction, it directly learns latent motion prediction.
The core idea is to train the model on "motion" rather than "visual details," allowing it to ignore interferences like smoke and reflections to model spatiotemporal action information. The authors highlight three key designs:
- guided masked prediction
- spatiotemporal self-distillation
- diversity regularization
Experimentally, SurgMotion achieved new SOTA results across multiple tasks:
- EgoSurgery workflow recognition: F1 improved by 14.6%
- PitVis: improved by 10.3%
- CholecT50: reached 39.54% mAP
Additionally, it covers tasks like skill assessment, polyp segmentation, and depth estimation. The post includes links to the paper, GitHub, Hugging Face, and an explainer report.
More from Research
- Anthropic masterclass spotlights how to build and observe AI agents — _jaydeepkarale · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21