ASIF: Coupled Data Turns Unimodal Models to Multimodal Without Training
Antonio Norelli, Marco Fumero, Valentino Maiorca, Luca Moschella, Emanuele Rodolà, Francesco Locatello
cs.LG, cs.AI, cs.CV
2022-10-05
ASIF aligns two frozen unimodal encoders with 1.6M image-text pairs and no training, reaching 60.9% ImageNet zero-shot versus 68.6% for CLIP trained on 400M pairs.
CLIP puts images and captions in one vector space and classifies by picking the closest class prompt. No classification layer is trained. The alignment is contrastive learning on enormous paired sets: 400M private pairs for the original model, 15M in the public release. LiT freezes a pretrained vision network and trains only the text encoder. The paired sets it reports still run from 10M to 901M.
Until the weights are public, a lab cannot use the model. A bad prediction is hard to trace to the pairs behind it. If a pair loses its license, removing its effect from a trained net is a machine-unlearning job. ASIF narrows the question. With two frozen unimodal encoders, a modest pile of image-text pairs, and no parameter update, how close does zero-shot classification get to CLIP?
Captions of nearby images tend to be nearby too. ASIF stores image-text pairs as anchors and writes each new sample as its similarities to those anchors. That vector is a relative representation: one coordinate per anchor, each a cosine similarity. Images are scored only against anchor images, texts only against anchor texts, and both vectors share a coordinate system. To pick a caption, treat the image vector as the relative representation of an ideal caption and take the nearest candidate. The bank behaves like a bilingual dictionary. Look the image up on the picture side, then match sentences against the captions of those same entries.
Anchors are the first 1.6M pairs of CC12M, web images with filtered alt-text. The vision encoder is either supervised DeiT-Base (ImageNet-1k, 768-d) or self-supervised DINO ViT-S/8 (ImageNet-21k, 384-d). The text encoder is a Sentence-Transformer trained on more than a billion internet sentences. All three stay frozen.
Raw similarities are noisy. Sparsification keeps the top k coordinates and zeros the rest. With n in the millions and k in the hundreds, small scores on unrelated anchors would drown the real neighbors. Surviving scores are raised to a power p at least 1, so the closest anchors dominate. The main runs use k = 800 and p = 8, chosen on the ImageNet validation set. Keeping k far below n is the choice that matters.
The embedding bank is the model. Writing or deleting one vector deploys a new version in seconds. At most k coordinates are nonzero, and each points at one pair, so a decision opens into a short list of training examples. Candidates can be sentences written at test time, so ASIF can stand in for CLIP on open-vocabulary classification. Inference scans the whole bank. Unoptimized, that pass is less than twice as slow as CLIP.
On ImageNet zero-shot, supervised ASIF scores 60.9% and the DINO variant 53.0%. The fair rows use a ViT-B/16-class encoder: CLIP on 400M pairs scores 68.6%, and LiT on 10M public CC12M pairs scores 66.9%. ASIF trails the first by about 8 points with 250 times fewer pairs, and the second by 6 points with about 6 times fewer pairs. Public CLIP on a curated 15M subset of YFCC100M manages only 31.3%. At 901M pairs with ViT-B/32, LiT scores 70.1% and the CLIP number reported by Zhai et al. is 50.6%. Those rows use the weaker backbone, so they are a poor basis for saying ASIF beats a 901M-pair CLIP. DINO is also smaller, 384-d against DeiT's 768-d, so the two ASIF rows are not a clean supervision comparison.
| Method | Pairs | ImageNet zero-shot |
| CLIP, ViT-B/16 | 400M private | 68.6% |
| CLIP | 15M public | 31.3% |
| LiT, ViT-B/16 | 10M public CC12M | 66.9% |
| CLIP, ViT-B/32 (Zhai et al.) | 901M private | 50.6% |
| LiT, ViT-B/32 | 901M private | 70.1% |
| ASIF, supervised DeiT | 1.6M public CC12M | 60.9% |
| ASIF, unsupervised DINO | 1.6M public CC12M | 53.0% |
Table 1 also lists CIFAR-100, Oxford Pets, and ImageNet-v2. In the extracted text those columns do not line up with all seven rows, so the cells are not repeated here. The paper places them in the same range as CLIP and LiT.
The first 10,000 pairs already reach 18% on ImageNet, against about 0.1% for a uniform guess over 1,000 classes. From a few hundred pairs to about a million, DeiT tiny, small, and base, plus smaller sentence encoders, keep improving. The curve does not flatten. The 1.6M cutoff is whatever fit on one Tesla T4.
On EuroSAT, ten satellite classes, unsupervised zero-shot ASIF scores 29.4% and CLIP 54.1%. After 100 EuroSAT pairs are encoded and stored, with no gradient step, accuracy is 82.2% ± 2.0. Product quantization, inverted indexes, and dropping anchors that never fire are discussed, without speed numbers.
Once the encoders exist, alignment is a data edit. DINO reaches 53.0% with no class labels on the vision side. The network is smaller than DeiT (384-d versus 768-d), so this is not a clean ablation, but labels are not the only source of the alignment. A revoked license is a deleted embedding. Labs that cannot assemble hundreds of millions of pairs get a baseline they can run.
As a classifier it is an incremental step. With enough pairs and a training budget to spare, ASIF still lands behind CLIP and LiT, and inference has to carry the full bank. The sharper question is how much of multimodal alignment was retrieval. A memory of embeddings improves as the memory grows, and it already sits close to contrastive training.
With abundant paired data, ASIF is behind CLIP and LiT. Relative representations stay high-dimensional after sparsification, so text-to-image generation is out of reach. The only measured task is zero-shot classification. Detection and retrieval are absent, and the datasets follow LiT.
The ImageNet cell does not bear much weight. Supervised DeiT was trained on ImageNet-1k and then scored on ImageNet. The prompts are new. The visual features are not. In Figure 6, a triumphal-arch query retrieves other arches and monuments, and the semantic gap is small. DINO at 53.0% rules out direct label leakage, and those features were still trained on ImageNet-21k images. The paper notes that LiT paired pretraining and test data the same way, so any gap is also a problem for CLIP and LiT. That hands the question back. It does not measure the distance. The paper asks for broader tasks and for that analysis. Neither is here.
k and p were picked on an ImageNet validation subset. Retuning on other datasets leaves accuracy roughly stable, and the headline table still uses hyperparameters from the same family of distributions. Captions can veto the encoders. If the neighbor texts are filenames or camera metadata, the class is not recoverable. CC12M alt-text is noisy web copy, which sits behind both the early 18% and the ceiling at 1.6M pairs. Whether ASIF would catch LiT at 15M or 400M pairs is not tested.