SLICEChat Prunes Visual Tokens During Encoding, Tops WSI-Bench With 1/5 Tokens

aykuterdemml · x · 2026-09-23

SLICEChat tackles gigapixel pathology slides by progressively pruning low-utility regions inside a hybrid Mamba-Transformer encoder, instead of compressing thousands of patch tokens after full encoding. The language-supervised, region-aware pruning keeps only 1/5 of visual tokens before multimodal fusion while retaining regions that matter for downstream reasoning. Results: 79.84% on TCGA SlideBench VQA, 59.09% on BCNB SlideBench VQA, best overall WSI-Bench metrics among evaluated models, with competitive memory and latency. Core idea: decide what matters before paying the full cost of encoding everything.

Original post →

More from Research

Research channel →