EvalSeal v2.2.0 adds tamper-evident eval receipts; LLM judge flipped 5/20 borderline cases
Fit_Fortune953 · reddit · 2026-09-28
EvalSeal is an open-source tool that makes LLM/agent eval results inspectable and reproducible. The v2.2.0 release adds: evalseal drift to classify why two runs disagree (evaluator drift, judge variance, target variance/change, or stable); evaluator fingerprinting across judge endpoint, model, temperature, topp, prompt template, rubric and dataset; preregister + gate --prereg to declare the eval contract before scores exist; and artifact binding for agent evals covering both tool ACLs and frozen tool responses.
A key finding: with an LLM judge, 5 of 20 borderline cases flipped verdicts across repeated runs, while numeric answer matching on 40 GSM8K cases flipped 0 — the instability came from the evaluator, not the model.
Documented limits: no local tool can prove someone didn't privately rerun until a good score appeared; requireciclaim is a claim, not proof; RFC 3161 support is experimental.
More from coding & agent
- 5 AI Agent Security Risks: More Autonomy Demands Tighter Controls — goyalshaliniuk · 2026-09-28
- Dev blasts Google AX's agent sandbox DX, offers 3-line alternative with Celesto — aniketmaurya · 2026-09-28
- Dev says $100-200 plans removed token limits; review cycles and human bandwidth are now the bottleneck — ezshine · 2026-09-28
- LangWatch benches open-source Jev alternatives on Jev's own benchmark — _rchaves_ · 2026-09-28
- Genomics MCP v0.1.0: 23 open-source tools let AI agents fetch genomic reads, variants and signal — Fair-Rain3366 · 2026-09-28
- Maybe don't let Muse run your Facebook Marketplace account — embedding-shape · 2026-09-28