NVIDIA Releases New Retrieval Embedding Model
kimmonismus · x · 2026-07-17
NVIDIA launched a new embedding model, Nemotron-3-Embed-8B, claiming it achieves an average NDCG@10 of 78.46 on RTEB and 75.45 on MMTEB Retrieval.
Official tests integrated it into a search agent powered by Nemotron 3 Ultra. Results show that more accurate retrieval surfaces relevant evidence earlier, reducing repeated searches, context checks, and reasoning turns, which lowers downstream token costs. Alongside the 8B version, NVIDIA released two 1B variants; the NVFP4 version optimized for Blackwell purportedly achieves up to 2× BF16 throughput while retaining over 99% retrieval quality. All three support 32K context, multilingual and code retrieval, enterprise RAG, and agent memory, with open weights, data, and recipes.
Related event: NVIDIA launches Nemotron-3-Embed retrieval models(11 posts)→
More from Infra
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11