faithgate: pytest for Prompts
ahumanbeingmars · reddit · 2026-07-05
Tired of the 'edit prompt, merge without testing' cycle, a developer created faithgate—essentially pytest for prompts: maintain a set of question/context/answer test cases, score faithfulness for a given prompt+model version, diff against baseline line by line, exit non-zero on regression, integrate with CI to block bad prompt changes at PR. Default uses Claude's built-in key as judge, metrics based on RAGAS; offline no-key mode: author publishes calibrated data (20 contradictions only 9 caught) and writes unit tests asserting that weakness. Failure means no pass: zero matches, ungraded, all wrong.
Related event: Faithgate Open-Sourced for RAG Fidelity Regression(2 posts)→
More from coding & agent
- Inspired by OpenAI's 10,000-agent run, dev open-sources a crowdsourced agent problem-solving platform — Benjaminsen · 2026-09-11
- Lucid: open-source Mac app keeps your laptop awake only while AI agents run — Pitiful_Hedgehog_600 · 2026-09-11
- banteg's snail project crowdsources AI agents to finish matching Snail Mail's 20 remaining functions — banteg · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11