Teaching Nemotron Greek: End-to-End RAG Adaptation and Benchmarking
KIEFERSA · hf · 2026-08-07
Addressing the absence of Modern Greek in major multilingual retrieval models, the authors present an end-to-end adaptation of NVIDIA's Nemotron retrieval stack, encompassing corpus mining, synthetic supervision, model training, and reader fine-tuning.
The study reveals that a parameter-free BM25 baseline surprisingly outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora. However, after fine-tuning on 65,773 Greek retrieval pairs, the Nemotron 1B embedder's nDCG@10 jumps from 0.362 to 0.835.
Furthermore, the authors LoRA-tune a Nemotron 30B-A3B mixture-of-experts reader for grounded generation, increasing answer correctness from 29.4% to 66.9% while significantly enhancing faithfulness and citation quality. The paper also introduces HERA, the first large-scale Greek benchmark for RAG systems.
More from Research
- Paradigm Shift in Continual Learning: From Parameter-Centric to System-Level Adaptation — CASIA · 2026-08-07
- TCFM: New Framework for Multilingual Text Embedding Adaptation via Flow Matching — LingoIITGN · 2026-08-07
- Curated List of Frontier Transformer-based SLAM Research — rsasaki0109 · 2026-08-07
- REI Labs Launches Adapt-1: A Pretraining-Free Architecture for Test-Time Learning — EnigmaFund · 2026-08-07
- OpenAI's Luna Aces ARC-AGI-1 at 90.7% with Massive 80% Cost Reduction — burny_tech · 2026-08-07
- ARC-AGI-3 Mechanics Clarified: Single Runs, No Shared State, Final Actions Scored — xeophon · 2026-08-07