EvalSeal v2.2.0 adds tamper-evident eval receipts; LLM judge flipped 5/20 borderline cases

Fit_Fortune953 · reddit · 2026-09-28

EvalSeal is an open-source tool that makes LLM/agent eval results inspectable and reproducible. The v2.2.0 release adds: evalseal drift to classify why two runs disagree (evaluator drift, judge variance, target variance/change, or stable); evaluator fingerprinting across judge endpoint, model, temperature, topp, prompt template, rubric and dataset; preregister + gate --prereg to declare the eval contract before scores exist; and artifact binding for agent evals covering both tool ACLs and frozen tool responses.

A key finding: with an LLM judge, 5 of 20 borderline cases flipped verdicts across repeated runs, while numeric answer matching on 40 GSM8K cases flipped 0 — the instability came from the evaluator, not the model.

Documented limits: no local tool can prove someone didn't privately rerun until a good score appeared; requireciclaim is a claim, not proof; RFC 3161 support is experimental.

Original post →

More from coding & agent

coding & agent channel →