SLICEChat Prunes Visual Tokens During Encoding, Tops WSI-Bench With 1/5 Tokens
aykuterdemml · x · 2026-09-23
SLICEChat tackles gigapixel pathology slides by progressively pruning low-utility regions inside a hybrid Mamba-Transformer encoder, instead of compressing thousands of patch tokens after full encoding. The language-supervised, region-aware pruning keeps only 1/5 of visual tokens before multimodal fusion while retaining regions that matter for downstream reasoning. Results: 79.84% on TCGA SlideBench VQA, 59.09% on BCNB SlideBench VQA, best overall WSI-Bench metrics among evaluated models, with competitive memory and latency. Core idea: decide what matters before paying the full cost of encoding everything.
More from Research
- CAIS benchmark: GPT-6 Astra automates 20.8% of remote work, up from 2.5% a year ago — scaling01 · 2026-09-23
- Eval scores reported to 5 significant digits despite 2-5 point error bars, researcher flags — dfrsrchtwts · 2026-09-23
- SkillSpec: intent-masked specification reasoning to catch semantic defects in agent skills — Buaa1 · 2026-09-23
- LLMs start overconfident, then swing underconfident when criticized, Nature MI paper finds — ValerioCapraro · 2026-09-23
- Nature MI paper: LLMs start overconfident, then swing underconfident when criticized — ValerioCapraro · 2026-09-23
- UK AISI paper: fixed-budget evals increasingly understate frontier LLM capability — evijit · 2026-09-23