Tsinghua's STAMPlus Solves MLLM Segmentation Trilemma with Single-Pass Inference
Tsinghua · hf · 2026-08-05
Tsinghua University introduces STAMPlus, a new architecture for Multimodal LLM (MLLM) segmentation that resolves the core trilemma between segmentation performance, dialogue ability, and inference speed.
- Core Innovation: It proposes Structured All-Mask Prediction, decoupling autoregressive dialogue from non-autoregressive mask prediction to jointly predict all targets in a single forward pass.
- Expanded Capabilities: A single unified checkpoint retains referring and reasoning capabilities while extending to open-vocabulary semantic, instance-aware, and remote-sensing small-target segmentation.
- Performance: While preserving general multimodal instruction following, it drastically reduces 12-category latency from 13.50s to 5.16s.
More from Multimodal
- Dual SAM3 + Seam Mask: A 4K Panoramic 3DGS Reconstruction Workflow — janusch_patas · 2026-08-05
- Using Hermes Desktop with ComfyUI: Letting AI Agents Auto-Fix Workflow Errors — Birdinhandandbush · 2026-08-05
- AI Demo Mimics Human Handwriting with Annotations and Streaming Charts — op7418 · 2026-08-05
- Handy ComfyUI Script: Audio Ping Notification for Job Completion — RPGstarDestroyer · 2026-08-05
- V2N: Multi-Task Visual Piano Transcription for Notes, Offsets, and Velocity — PianoVAM · 2026-08-05
- Kyutai Launches Muscriptor: High-Precision Audio-to-MIDI Model — huggingface · 2026-08-05