EvalSeal v1.5.0: open-source reproducibility receipts for LLM evals
Fit_Fortune953 · reddit · 2026-09-21
A developer released EvalSeal v1.5.0, an open-source tool that makes LLM eval results trustworthy by running cases multiple times, measuring per-case instability, capturing provenance, and sealing results into a tamper-evident ledger.
Features
- Per-case verdict strips and evaluator fingerprinting
- Signed ledger heads, CI gates, drift comparison via evalseal diff
- Deterministic HTML receipts attachable to PRs; replayable examples without API keys
Key finding: with an LLM judge, 5 of 20 borderline cases flipped verdicts across repeated runs; numeric answer matching on 40 GSM8K cases flipped 0. Same model family — the instability came from the evaluator, not the target model.
Related event: EvalSeal Open-Sourced to Bring Trust Receipts to LLM Evaluations(2 posts)→
More from coding & agent
- hearim: open-source gateway turns ordinary local LLMs into a Jev-compatible decision API — ziozzang0 · 2026-09-21
- Tencent open-sources T-Mem, an associative-recall memory architecture for long-term AI memory — aigclink · 2026-09-21
- Indie founder building his second $1M business launches AI-employee platform — marclou · 2026-09-21
- Jev-as-a-Judge: LangChain says typed evaluators can cut RL verification costs by orders of magnitude — NandoDF · 2026-09-21
- What failure cases must an LLM gateway pass before automatic failover is safe? — Rama_Surasani_ · 2026-09-21
- HarnessRouter open-sources one API to run Codex, Claude Code and 14 agent harnesses — gaganghotra_ · 2026-09-21