AutoRubric, a Rubric-Science-Backed LLM Evaluation Tool, Heads to COLM 2026

deliprao · x · 2026-09-21

AutoRubric—combining rubric science best practices with LLM-as-a-Judge research for structured evaluation of non-verifiable domains—has been accepted to COLM 2026. It lets you define weighted criteria (including negative scores for wrong claims), validate against human labels, and iterate via built-in meta-evaluation; an open-source Python library, cookbook, and example grading NMC vs LFP battery answers with gpt-5.1-mini are available.

Related event: Researcher Alleges TypeSafe's Jev Mirrors His AutoRubric Paper(3 posts)→

Original post →

More from coding & agent

coding & agent channel →