NVIDIA Releases Multilingual Embedding Models

tomaarsen · x · 2026-07-17

NVIDIA has released a new embedding solution natively integrated into Sentence Transformers. Users can directly load nvidia/Nemotron-3-Embed-8B-BF16 to encode queries and documents separately.

The post notes support for an NVFP4 variant and vLLM. It preserves prompts and normalization metadata while automatically handling query/passage prefixes. The author considers this a major release for multilingual RAG, showing leading retrieval accuracy, and suggests the 1B version is better suited for budget-conscious teams.

Related event: NVIDIA launches Nemotron-3-Embed retrieval models(11 posts)→

Original post →

More from Infra

Infra channel →