Open-source RAGnarok-AI launches human-annotation study: can you trust LLM judges?
Ok-Swim9349 · reddit · 2026-09-05
The maintainer of RAGnarok-AI, an open-source local-first RAG evaluation framework, is running a study on whether automated LLM-judge evaluations can be trusted. The framework scores retrieval relevance, faithfulness, answer relevance, and completeness; now he's building a human-annotated benchmark against those scores and recruiting developers for 10-15 anonymous annotation cases (30-45 min, no RAG expertise needed). Covering docs from Docker, Python, FastAPI, and Kubernetes, the study tests judge reliability, discrimination between good and degraded RAG systems, and reproducibility — all methodology versioned in public.
More from Research
- RL Hill-Climbing Creates Model Spikes, So Judge AI With Multiple Models — dejavucoder · 2026-09-05
- Study: Generative AI Has Already Cost Philippines' BPO Industry ~250,000 Jobs — sebkrier · 2026-09-05
- NTU Singapore to Host GDL2026 Conference on Geometry, Dynamics, and Learning, Sept 28-30 — FrnkNlsn · 2026-09-05
- Flow Reasoning Models hit 99.5% on Sudoku-Extreme with 44× fewer inference FLOPs — mark_k · 2026-09-05
- Conditional independence limits parallel token generation in masked diffusion models — alec_helbling · 2026-09-05
- Seven-minute chatbot conversation beats fact sheets at reducing conspiracy beliefs — The Decoder · 2026-09-05