Scaling experiment: 10k pairs, 75 mins training, 1.2 points short of full run
tomaarsen · x · 2026-08-26
Experiments on an RTX 3090 show that training with only 100k pairs (10% of data) for 75 minutes yields NDCG@10 only 1.2 points lower than a full million-pair run, with most gains occurring in the first hour.
Comparisons reveal that Late Interaction models significantly outperform Dense models, beating the much larger Qwen3-Embedding-4B by 13 points. Notably, Qwen3's 8B version scores lower than its 4B counterpart.
More from Research
- PrimeIntellect verifiers v0.3.1: Model Interception and Persistent ACP Sessions — xeophon · 2026-08-27
- RAG Isn't Dead: Navigating Retrieval vs. Agentic Search — hugobowne · 2026-08-27
- AI audit not infallible: Refine missed a known lemma error in paper — littmath · 2026-08-27
- AI audit of 19 papers: 97.7% of comments identified real issues — littmath · 2026-08-27
- Mathematician's AI-generated errata are "slop" but mostly correct — littmath · 2026-08-27
- Mathematician Litt: Paper Errors Mostly From Badly Propagated Edits, Not Deep Flaws — littmath · 2026-08-27