NVIDIA open-sources NV-Reason-CT, a native 3D vision-language model for CT scans

NVIDIA Developer · youtube · 2026-10-09

NVIDIA released NV-Reason-CT, an open vision-language model that natively reads chest and abdominal CT volumes in 3D and reasons step by step to produce findings. A Primus 3D ViT encoder feeds Qwen3.5-4B end to end, passing all 13,824 visual tokens per 192×192×192 crop with 3D position tags — unlike 2D-oriented VLMs that sample slices or compress volumes. Model, demo, fine-tuning scripts (SFT/GRPO), and paper are public.

Original post →

More from Models

Models channel →