Open-source RAGnarok-AI launches human-annotation study: can you trust LLM judges?

Ok-Swim9349 · reddit · 2026-09-05

The maintainer of RAGnarok-AI, an open-source local-first RAG evaluation framework, is running a study on whether automated LLM-judge evaluations can be trusted. The framework scores retrieval relevance, faithfulness, answer relevance, and completeness; now he's building a human-annotated benchmark against those scores and recruiting developers for 10-15 anonymous annotation cases (30-45 min, no RAG expertise needed). Covering docs from Docker, Python, FastAPI, and Kubernetes, the study tests judge reliability, discrimination between good and degraded RAG systems, and reproducibility — all methodology versioned in public.

Original post →

More from Research

Research channel →