EvalSeal: open-source tool shows LLM judges flip verdicts on 5 of 20 borderline eval cases
Fit_Fortune953 · reddit · 2026-09-18
- EvalSeal (v0.3.0) is a small open-source tool for making LLM eval results trustworthy: it runs each case multiple times, measures flip rates, captures provenance, and seals results into a tamper-evident ledger.
- Author's findings: with an LLM judge, 5 of 20 borderline cases flipped verdicts across repeated runs; with exact numeric matching on 40 GSM8K cases, 0 flipped.
- Key takeaway: same model family — the instability came from the judge, not the target model, so a single eval score isn't enough.
- v1 roadmap: signing, suite files, benchmark-oriented scoring. Feedback wanted from folks working on evals, CI gates, and agent reliability.
More from coding & agent
- Pydantic AI agents now run on TypeSafe's Jev with per-field confidence outputs — Paimaamu · 2026-09-18
- Demo: GPT-6 Astra edits video in DaVinci via MCP at impressive speed — toolstelegraph · 2026-09-18
- After a day with Jev: a blazing-fast classifier, not a GPT replacement — jiayuan_jy · 2026-09-18
- Wiring Claude Code, Codex and Cursor to one shared MCP memory server — Asly97 · 2026-09-18
- Running Claude Code, Codex and Cursor on One Project: How I Finally Stopped Losing Context — Asly97 · 2026-09-18
- Multica hits 50k GitHub stars, ships cloud agent integration with Qoder — jiayuan_jy · 2026-09-18