SurgMotion: World's First Billion-Parameter Surgical Video Foundation Model Released
机器之心 · wechat · 2026-07-08
The CAS Hong Kong Innovation Research Institute, along with the Institute of Automation CAS, The Chinese University of Hong Kong, Technical University of Munich, and several top hospitals, have released SurgMotion, the world's first native billion-parameter surgical video foundation model. Trained on the self-built SurgMotion-15M dataset—which includes 3658 hours of surgical video, 15 million frames, covering 13 anatomical regions and 50 data sources—it is currently the world's largest pre-training dataset for surgical videos. SurgMotion adopts the V-JEPA architecture and introduces three core technologies, including motion-guided latent space masked prediction and spatiotemporal affinity self-distillation. It achieved an average performance improvement of 16.5% in dynamic understanding and an average error reduction of 2.9% in static tasks across a benchmark covering 17 core surgical tasks. Within three months of its release, the model ranked first in downloads for surgical video foundation models on HuggingFace. Nearly 40 top institutions across 14 countries on 5 continents have applied for access, including Intuitive Surgical, KARL STORZ, and Duke University.
More from Research
- Noahpinion quotes Chollet: intelligence may hit a hard ceiling — binarybits · 2026-07-27
- Paper argues graph topology can become the core operating system for AI agents — theomitsa · 2026-07-27
- A question probes how multi-agent branching scales against compute budget and model size — iskander · 2026-07-27
- NUS builds a soft force sensor that drives actuators without electronics or power — CurieuxExplorer · 2026-07-27
- Chelsea Finn says robot RL is bottlenecked by physical rollout cost, not algorithms — ycombinator · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27