Delip Rao launches Autorubric at COLM 2026 to unify rubric-based LLM evaluation
deliprao · x · 2026-10-09
At COLM 2026, Delip Rao's team presented "Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks".
The problem
- Rubric-based LLM-as-a-Judge is now standard for non-verifiable tasks, but techniques are fragmented across literature with inconsistent implementations.
- A rubric is not "just a prompt" — rubric judges can silently fail in multiple ways.
- Without shared infrastructure, researchers pay a repeated "reinvention tax", and many well-known papers shipped implementation errors whose fixes already existed in prior work.
What Autorubric does
- Defines rubrics as generally and rigorously as possible.
- Ships state-of-the-art measurement features: ensembling, calibration, reliability methods.
- Gives eval and reward-modeling researchers shared infrastructure with best practices built in for efficiency and correctness.
More from Research
- Could Looped Models Resist Distillation Attacks by Reasoning in Latent Space? — moyix · 2026-10-09
- LLM-as-a-Verifier: Weaker Model Verifies Stronger One, Hits 69.2% SOTA on Terminal-Bench 4 — Azaliamirh · 2026-10-09
- Salesforce's SRD distills hindsight into foresight, lifting 2B agent success from 0% to 60.6% — Salesforce · 2026-10-09
- Claude claims discovery of new binary red dwarf pair ~500 light-years away — nitarshan · 2026-10-09
- Prompt Tuning Is Forgotten Lore — Are We Massively Underusing Finetuned Tokens? — cephaloform · 2026-10-09
- Ricardo Baeza-Yates Lecture: When Will ML Evaluation Stop Fooling Itself? — PolarBearby · 2026-10-09