faithgate: Regression Testing Gate for RAG Faithfulness
ahumanbeingmars · reddit · 2026-07-06
After a prompt change silently caused their RAG bot to hallucinate without errors for days, the author open-sourced faithgate. It acts as a regression gate for faithfulness: maintaining a set of question/context/answer test cases to score every prompt or model version. If grounding drops below the baseline, CI fails. It uses the RAGAS Faithfulness metric under the hood with Claude as the default judge. Testing on a 20-document corpus with three planted hallucination types (date swaps, entity swaps, and baseless splicing) caught them all, dropping scores from 1.00 to 0.29, 0.12, and 0.20. The author notes the offline keyless judge mode is weaker, catching only 9/20 unfaithful answers in 40 manual labels, and includes unit tests asserting this number.
Related event: Faithgate Open-Sourced for RAG Fidelity Regression(2 posts)→
More from coding & agent
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- ARRM targets silent economic regressions in AI agents that functional tests miss — Beautiful_Belt_601 · 2026-09-11
- Dev builds browser 3D pizza delivery game with Claude: physics, GPS pathfinding, traffic AI — vinishkapoor · 2026-09-11
- Build X Carousel Posts from One Wide Image: A Splitter Tool Plus YouMind Skill Workflow — sujingshen · 2026-09-11
- "Anyone still coding the old way?" The joke capturing post-AI programming culture — lxfater · 2026-09-11