lm-eval-ledger: an open-source harness to browse and compare per-question model answers

jayminban · reddit · 2026-09-07

Developer jayminban released lm-eval-ledger, an open-source LLM benchmark harness fixing the pain that existing tools only hand you headline scores while answers are buried in JSONL/Parquet dumps.

Live demo on Hugging Face Spaces, code on GitHub, feedback and PRs welcome.

Original post →

More from coding & agent

coding & agent channel →