SurgMotion: World's First Billion-Parameter Surgical Video Foundation Model Released

机器之心 · wechat · 2026-07-08

The CAS Hong Kong Innovation Research Institute, along with the Institute of Automation CAS, The Chinese University of Hong Kong, Technical University of Munich, and several top hospitals, have released SurgMotion, the world's first native billion-parameter surgical video foundation model. Trained on the self-built SurgMotion-15M dataset—which includes 3658 hours of surgical video, 15 million frames, covering 13 anatomical regions and 50 data sources—it is currently the world's largest pre-training dataset for surgical videos. SurgMotion adopts the V-JEPA architecture and introduces three core technologies, including motion-guided latent space masked prediction and spatiotemporal affinity self-distillation. It achieved an average performance improvement of 16.5% in dynamic understanding and an average error reduction of 2.9% in static tasks across a benchmark covering 17 core surgical tasks. Within three months of its release, the model ranked first in downloads for surgical video foundation models on HuggingFace. Nearly 40 top institutions across 14 countries on 5 continents have applied for access, including Intuitive Surgical, KARL STORZ, and Duke University.

Original post →

More from Research

Research channel →