NVIDIA open-sources NV-Reason-CT: 3D CT VLM hits 0.614 F1 on CT-RATE with no classification head

NV-Reason-CT: 3D Visual Language Model for CT Analysis

Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Zongwei Zhou, Wenxuan Li, Marc Edgar, Yufan He, Pengfei Guo, Daguang Xu

cs.CV, cs.AI

2026-09-23

NVIDIA couples a native 3D ViT with Qwen3.5-4B, feeding all 13,824 CT visual tokens with explicit 3D coordinates into the LLM. CT-RATE macro-F1 0.614 with no classification head; report F1 0.592. Open-sourced.

What problem this solves

CT is inherently volumetric: one exam spans hundreds of slices, and the diagnosis hangs on morphology, extent, laterality, and anatomical distribution across them. Existing CT vision-language models compromise at exactly this point. CT-CHAT attention-pools visual features into 256 tokens before Llama sees them; Jolia pools an entire abdominal CT into a single token; M3D-LaMed, RadFM, and VoxelFM pass through Q-Former or Perceiver resamplers. The vision tower may encode 3D geometry internally, but the language decoder receives a flattened 1D sequence with no explicit depth-height-width grid. Telling a left renal lesion from a right one, or describing craniocaudal extent, depends on the information dropped at that interface.

Supervision has a second gap. Radiology reports state conclusions, not the path from image to conclusion. SFT on reports alone teaches a model what to say, not how the finding was reached.

Method

The central architectural decision: no token merging, and coordinates travel with the tokens.

Training data covers about 550,000 instruction examples from 70,111 CTs (CT-RATE 47k, internal NIH 16k, CancerVerse 23k). The key ingredient is recorded narrations from radiologists reading cases as usual: transcribed for direct supervision, then used as exemplars to rewrite source reports into synthetic narrated reasoning with observations, differentials, and uncertainty, with the source report constraining clinical content. The corpus adds structured reports (9 chest sections, 13 abdominal), binary abnormality QA, localization and severity QA, multi-turn follow-ups, and refusal examples for non-medical requests and invalid inputs.

Training is two-stage: end-to-end SFT of encoder, projector, and LLM together (with component learning rates deliberately split, 0.1× for the encoder and 5× for the projector), then GRPO with fully verifiable rewards: 2.0 × finding-set F1, 0.5 × report-structure conformity, and a length penalty outside a 180-900 token band, with a machine-readable <answer> tag required at the end of each report. RL needs only abnormality labels, not an expert reasoning trace per volume. Trained on 16 nodes with 128 H100 GPUs.

Results

BenchmarkMetricNV-Reason-CTBest baseline
CT-RATE classification (18 labels)macro-F10.614VoxelFM 0.581 (trained head)
CT-RATE classificationmacro-AUROC0.871VoxelFM 0.870
CT-RATE report generationreport-derived macro-F10.592CT-AGRG 0.501
Preliminary reader studyavg interpretation + reporting time50% reductionunassisted

The 0.614 comes from generated Yes/No text, with no task-specific classification head; the comparators VoxelFM (0.581), CT-SSG (0.572), and Pillar-0 (0.544) all rely on separately trained classifiers. A generative interface outscoring every supervised classifier in this table is the notable part. Switching from per-abnormality questions to a single prompt listing all findings drops F1 only from 0.614 to 0.610, at an order of magnitude less inference cost. Report generation beats CT-AGRG by 0.09; CT-AGRG is effectively a classifier gating sentence-by-sentence generation. On external benchmarks it ranks first with RAD-ChestCT F1 0.766 and Merlin abdomen 0.531, under slightly different protocols.

Why it matters

Weights are on HuggingFace and the full SFT plus GRPO training code is on GitHub. For medical multimodal teams this is a rare reference implementation where data construction, training scripts, and final weights are all public, at a deployable 4B scale. The all-tokens-plus-explicit-3D-coordinates interface is portable to other Qwen-based models regardless of scale. Reviewable observations, differential diagnoses, and stated uncertainty sit closer to actual radiology workflow than a classification score; if the halved reading time holds up, the value is immediate.

Limitations

Stated by the authors: the cropping heuristic depends on visible lungs and can mislocalize when lung landmarks are absent; generated reasoning is framed as a reviewable clinical explanation, with no assumption that it faithfully reflects internal computation. From reading the paper: the core design of unmerged tokens with coordinates has no controlled ablation, so its contribution cannot be separated from the 550k-example data effort. A single 2mm, 384mm crop means a few-millimeter nodule occupies less than one 8³ patch token, a real concern for small-lesion detection in a benchmark dominated by lung nodules. Training includes 15,991 non-public internal NIH CTs, so full reproduction is impossible. Comparison protocols are messy across v1/v2 manifests, reconstruction- versus study-level units, and thresholds; CT-CLIP's originally reported 0.7069 F1 falls to 0.398 under the unified protocol. The AnyMC3D ten-model ensemble with calibration reaches 0.646, above the single-model 0.614. The reader study is small, timing is self-reported, and the paper says "associated with" rather than claiming causation.

Terms

Source

What people are saying

Related papers

All paper explainers