Modular Agent Outperforms End-to-End VLMs in CT Spatial Reasoning

Simon Vincent Abel · hf · 2026-08-27

Presents a modular medical imaging agent for verifying spatial relations in CT scans. By decomposing the task into parsing, anatomical localization, and geometric rules, this approach outperforms end-to-end vision-language models on CT spatial reasoning benchmarks, offering improved reliability and auditability.

Original post →

More from Research

Research channel →