NVIDIA Releases NV-Reason-CT: LLM Reasoning Over 13,824 Visual Tokens Per CT Scan
techwith_ram · x · 2026-09-27
NVIDIA has released NV-Reason-CT, a 3D medical AI model that lets an LLM reason over a single CT scan, feeding in 13,824 visual tokens at once.
Architecture: a Qwen3.5-4B language backbone paired with a Primus 3D vision encoder initialized from COLIPRI. Input is one-channel chest or abdominal CT (.nii / .nii.gz), preprocessed with LPS orientation, 2mm isotropic resampling, and an anatomy-aware 192×192×192 voxel crop. A 24×24×24 patch grid of non-overlapping 8×8×8 patches produces the tokens, which are passed to the LLM without spatial merging, with depth/height/width encoded via 3D MRoPE.
Available on GitHub and Hugging Face.
More from Models
- Musk confirms Grok 'upgrades' as users notice dramatic speed boost — elonmusk · 2026-09-28
- "System 2 models built the brain, but System 1 is building the nervous system" — ai · 2026-09-28
- TeleOCR Trends on Hugging Face: A Qwen2.5-VL-Based Chinese Document OCR Model — XingChen-AGI · 2026-09-28
- Kaggle Game Arena: Evaluating LLMs via Head-to-Head Chess, Poker, and Werewolf — kaggle · 2026-09-28
- Perplexity CEO: still using sol 6 for knowledge work — cheap, fast, great compaction — gabriel1 · 2026-09-28
- NerfBench's First Results Find No Nerf: Claude Opus 5.5 Dips Just 0.8% vs Launch — alejandroll10 · 2026-09-28