Teaching Nemotron Greek: End-to-End RAG Adaptation and Benchmarking

KIEFERSA · hf · 2026-08-07

Addressing the absence of Modern Greek in major multilingual retrieval models, the authors present an end-to-end adaptation of NVIDIA's Nemotron retrieval stack, encompassing corpus mining, synthetic supervision, model training, and reader fine-tuning.

The study reveals that a parameter-free BM25 baseline surprisingly outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora. However, after fine-tuning on 65,773 Greek retrieval pairs, the Nemotron 1B embedder's nDCG@10 jumps from 0.362 to 0.835.

Furthermore, the authors LoRA-tune a Nemotron 30B-A3B mixture-of-experts reader for grounded generation, increasing answer correctness from 29.4% to 66.9% while significantly enhancing faithfulness and citation quality. The paper also introduces HERA, the first large-scale Greek benchmark for RAG systems.

Original post →

More from Research

Research channel →