EvalSeal v1.5.0: open-source reproducibility receipts for LLM evals

Fit_Fortune953 · reddit · 2026-09-21

A developer released EvalSeal v1.5.0, an open-source tool that makes LLM eval results trustworthy by running cases multiple times, measuring per-case instability, capturing provenance, and sealing results into a tamper-evident ledger.

Features

Key finding: with an LLM judge, 5 of 20 borderline cases flipped verdicts across repeated runs; numeric answer matching on 40 GSM8K cases flipped 0. Same model family — the instability came from the evaluator, not the target model.

Related event: EvalSeal Open-Sourced to Bring Trust Receipts to LLM Evaluations(2 posts)→

Original post →

More from coding & agent

coding & agent channel →