Scaling experiment: 10k pairs, 75 mins training, 1.2 points short of full run

tomaarsen · x · 2026-08-26

Experiments on an RTX 3090 show that training with only 100k pairs (10% of data) for 75 minutes yields NDCG@10 only 1.2 points lower than a full million-pair run, with most gains occurring in the first hour.

Comparisons reveal that Late Interaction models significantly outperform Dense models, beating the much larger Qwen3-Embedding-4B by 13 points. Notably, Qwen3's 8B version scores lower than its 4B counterpart.

Related event: Finetuning ColBERT on a single RTX 3090 beats general-purpose retrievers in medical search(13 posts)→

Original post →

More from Research

Research channel →