NVIDIA Releases NV-Reason-CT: LLM Reasoning Over 13,824 Visual Tokens Per CT Scan

techwith_ram · x · 2026-09-27

NVIDIA has released NV-Reason-CT, a 3D medical AI model that lets an LLM reason over a single CT scan, feeding in 13,824 visual tokens at once.

Architecture: a Qwen3.5-4B language backbone paired with a Primus 3D vision encoder initialized from COLIPRI. Input is one-channel chest or abdominal CT (.nii / .nii.gz), preprocessed with LPS orientation, 2mm isotropic resampling, and an anatomy-aware 192×192×192 voxel crop. A 24×24×24 patch grid of non-overlapping 8×8×8 patches produces the tokens, which are passed to the LLM without spatial merging, with depth/height/width encoded via 3D MRoPE.

Available on GitHub and Hugging Face.

Original post →

More from Models

Models channel →