Ten vision models tie at r=0.67, then fall to r=0.45 under neural control

2026-10-10

Ten vision models averaged r=0.67 on natural images in five macaques, then fell to r=0.45 driving the same 25 sites with 27,500 axis-aligned images. Only two robust models held up.

What problem this solves

Encoding leaderboards for visual cortex have piled up. On Brain-Score and large fMRI benchmarks, networks with very different architectures and training objectives predict cortical responses about equally well. That invites a strong inference: the models have converged on one brain-aligned parameterization of natural images.

The inference rests on observational regression. Natural images are densely covariant, so many incompatible feature sets can fit the same firing rates to a similar correlation. The name for that pile-up is the Rashomon effect: one dataset, many explanations that all look right. Recovering a neuron's actual tuning needs an intervention. Change the image the way the model says the neuron should care about, then see whether firing follows.

Method

The experiment is a four-phase loop. Five macaques, 25 sites, five per animal, covering V3, V4, central and anterior inferotemporal cortex (cIT, aIT), and the superior temporal sulcus (STS). The two IT animals used chronic arrays; the other three used Neuropixels. Sites were picked for calibration reliability, and a candidate was dropped if its tuning correlated above 0.9 with a site already chosen. Mean calibration noise ceiling was r=0.88.

The calibration set has 969 images: 649 NSD natural scenes, 259 segmented objects on white, 61 functional-localizer images. 774 fit the readout. 195 select the layer and supply the encoding score.

Ten models were chosen to spread the usual contrasts while sharing backbones where possible, so a difference can be pinned on a training pressure. The convolutional set is ResNet50, language-contrastive ResNet50-CLIP, self-supervised ResNet50-DINO, and SEER-trained RegNetY-640. The transformers are DINOv2 (ViT-B/14, with registers), SigLIP2 (ViT-B/16), and RADIO v2.5-B, which distills several foundation models into one backbone. Two lines use adversarial training: ResNet50-Robust, and CLIPAG, an adversarially fine-tuned CLIP ViT-B/32. The tenth is the AlexNet checkpoint that was leading Brain-Score's V1 benchmark through that API. Its weights barely match a properly trained AlexNet, and its classification accuracy sits at an untrained level. In this study it is the untrained baseline.

Each site and each model gets its own encoding fit. Layer activations are reduced to 750 principal components, then ridge regression, with the layer chosen by validation R². The weight vector is the encoding axis: move along it and the model predicts that the site's firing will rise or fall. PCA and ridge regression are both linear, so the map from pixels to predicted firing is differentiable end to end.

Stimuli come from axis-aligned feature accentuation. Treat the axis as a ruler. Starting from a seed image, gradient updates run in the Fourier domain, with a penalty on high spatial frequencies plus random crops and noise, pushing predicted firing to 11 levels. The span runs from 25% below the site's natural-image response range to 50% above it. Regularization is not hand-tuned. An automatic ladder picks a setting per model-site pair until synthesis success clears a threshold. Ten seeds, 25 sites, 10 models, 11 levels: 27,500 images by design. At a 5% tolerance, 94.5% of images hit their target level. Those images go back to the same sites.

Control is scored two ways: Pearson r between predicted and measured firing, and the slope of measured on predicted. A slope of 1 means the levels and the spikes line up.

Results

On natural images the ten models average validation r=0.67, with a standard deviation of 0.07 across model means. The nine trained models have a standard deviation of only 0.01, and their validation r sits between 0.67 and 0.71. The untrained baseline is r=0.47. The procedure can tell trained from untrained. It cannot separate the leaders.

On each model's own accentuated images, mean control r falls to 0.45. The standard deviation across model means rises to 0.15, and across the 250 model-site pairs it is 0.27. A repeated-measures ANOVA gives F(9, 216)=25.5, p=1.4e-29. The model effect is significant inside each of the five areas. Rank agreement across areas averages Spearman 0.69, and aIT versus cIT is 0.94.

One anterior-IT site with ResNet50 shows the gap in one place. Calibration validation r=0.87, and a retest of held-out natural images days later is r=0.95. Along that same axis, control r is 0.28 and the slope is 0.05. The sweep is nearly flat. The noise ceiling on repeated presentations of those images is r=0.50, so a noisy session does not explain the miss. At the other end, CLIPAG reaches r=0.91 and slope 0.81 at one V3/V4 site. Mean control r across the 25 sites runs from 0.09 to 0.74.

The two models that stand out are the adversarially trained ones, ResNet50-Robust and CLIPAG. Convolution versus transformer, self-supervised versus supervised, language-aligned versus vision-only: none of those groupings separates control. A controversial set that pushes ResNet50 and its robust twin in opposite directions on the same image leaves unique firing-aligned variance only on the robust side.

Adversarial sensitivity does not explain the gradient inside the leaderboard. Across all models, sensitivity and control slope correlate at site-level r=-0.51 (more sensitive, worse control) and model-level r=0.89. Drop the two robust models and the site-level figure falls to 0.05, model-level to 0.41, permutation p=0.36. The steadier predictor is spectral concentration of the input gradient: backpropagate the encoding axis to pixels and ask whether energy sits on low-frequency structure such as object boundaries, or spreads as high-frequency noise. Across all models, site-level r=0.64 and model-level r=0.90. After dropping the robust models, site-level r is still 0.26 (p=0.0006) and model-level r is 0.93 (p=0.0016). The measure does not need synthesized images. It can be computed before any control session. Held out to unseen models, sites, and animals, R² is about 0.42 to 0.45.

ComparisonMetricResult
Natural-image encodingMean r, 10 models0.67±0.07; SD=0.01 among 9 trained
Untrained baselineEncoding r0.47
Control on accentuated imagesMean r0.45±0.15; SD=0.27 across 250 pairs
Model differencesRepeated-measures ANOVAF(9,216)=25.5, p=1.4e-29
aIT × ResNet50Natural / controlval 0.87, retest 0.95; control r=0.28, slope 0.05
V3/V4 × CLIPAGControlr=0.91, slope 0.81
Robust models removedSpectral concentration vs sensitivitysite-level r=0.26 vs 0.05

Personalized sets of 110 images, the sweep built along that model-site axis, separate models with a standard deviation of 0.11. Pooling 5,500 accentuated images per monkey shrinks that to 0.04. Removing each model's own images shrinks it to 0.02. Robust models stay slightly ahead even in the pool, mean r 0.48 versus 0.40, paired t(24)=5.3, p=2e-5. They remain top-ranked after every image they generated is removed.

Against natural images, the strongest available comparison takes the 220 ImageNet validation images, out of 50,000, that maximize predicted disagreement for each model pair. At one anterior-IT site, ResNet50 versus ResNet50-Robust disagrees by 0.60 standardized units on that ImageNet subset and by 1.30 on the accentuated set, about 2.1 times. Accentuated images win in 90.5% of 1,125 channel-model-pair combinations, and at 24 of 25 sites, Wilcoxon p=1.2e-7.

Why it matters

For anyone fitting encoding models or reading brain-alignment leaderboards, linear predictivity on natural images no longer separates the top networks. The test here is different: the model names an axis, and the neuron either follows it or does not. A pre-screen does not require a new animal session. Compute how concentrated the encoding axis's pixel gradient is in spatial frequency. Adversarial training is currently the most effective way to produce that structure. The degree of robustness itself is not a continuous predictor of control among conventionally trained models.

For feature visualization, closed-loop stimulation, and prosthetic-style control, one number is enough to change a default. An axis with r=0.87 on natural images can be flat when asked to drive firing. Control experiments should try robust backbones first, or models whose gradient spectra are more concentrated.

The untrained AlexNet sitting atop a Brain-Score V1 ranking is a side finding, and it is a practical one. The leaderboard scores variance a linear readout can extract. Whether the checkpoint was actually trained is a separate check.

This is a change in the evaluation standard. Another tenth of correlation on natural images, from a new architecture, does not by itself mean the tuning is more accurate.

Limitations

The calibration encoding scores are validation scores after layer selection, not scores on images that never entered that choice. Control sessions started 4 to 8 days later, and cross-session retest r on natural images fell to the range 0.51 to 0.54. Models still cluster tightly on that retest while control scores spread, so the ranking is hard to blame on session drift alone. How much of the absolute drop is the gap between days is not fully partitioned.

Leave-one-out R² for the spectral predictor is only 0.42 to 0.45. The paper is explicit that a cleaner gradient should not be treated as the mechanism of control. It may be a proxy for some hierarchical alignment that is not yet isolated. Spatial-frequency bias and adversarial robustness can also come apart.

Only the on-axis direction was tested. Whether the neuron is actually invariant along directions the model calls a null space is open. Twenty-five sites in five animals, with the two IT implants aimed at fMRI-localized face patches, is a thin sample of ventral cortex.

High-frequency penalties and image augmentations during synthesis could favor some models. The paper's reply has two parts. Robust models still predict better on images other models synthesized. The regularizer was built to make non-robust models produce changes a person can see. That reply is coherent, and it is not a same-axis, two-synthesizer ablation. Procedure and model stay tied together. The authors also note that an internet-scale image search might close some of the gap with synthesized images. The comparison here stops at the best subset of 50,000 ImageNet validation images.

Terms

Source

What people are saying

All paper explainers