SAM-MT: Real-Time Multi-Object Video Segmentation
FudanCVL · hf · 2026-07-11
This work proposes SAM-MT for real-time interactive multi-object video segmentation. The authors point out that while existing methods perform well on single targets, extending them to multiple targets usually requires repeating the entire single-target pipeline for each object. This tanks the frame rate, with latency becoming uncontrollable as targets increase.
Method
- Modifies Segment Anything 2 (SAM2) into a multi-object interaction framework
- Uses explicit queries to represent different targets while retaining global context in parallel
- Reduces interference between targets and maintains identity distinction via decoupled masked attention
- Combines sparse memory for temporal stability, adding occlusion handling and overlap prevention strategies
Results
- Decouples latency from the number of targets
- Maintains >36 FPS even with 10 targets
- Performance approaches single-target baselines while preserving SAM2's video segmentation quality
More from Multimodal
- One prompt, full UGC ad: Kling MCP turns a product idea into ready-to-post video — SimplyAnnisa · 2026-09-11
- A Seedance 2.5 quick-start prompt with GPT Image 2.5 hacks — techhalla · 2026-09-11
- TestingCatalog's Daily AI Brief adds email editions, dishing Meta Muse and GPT-Live-1 rumors — testingcatalog · 2026-09-11
- Reddit user shares WIP AI-generated dark fantasy short film 'Wanderers' — DaWid_Shapiro · 2026-09-11
- Creator makes 2D electro-pop anime music video with just a prompt using MiniMax H3 — Hailuo_AI · 2026-09-11
- Street View to driving footage: GPT Astra fetches images, MiniMax H3 turns them into dashcam video — Hailuo_AI · 2026-09-11