Autorubric at COLM 2026: a unifying framework for rubric-based LLM evaluation
deliprao · x · 2026-10-09
Delip Rao's team is presenting Autorubric at COLM 2026, a unifying framework for rubric-based LLM evaluation on non-verifiable tasks, aimed at practitioners doing evals and reward modeling.
The idea is to make open-ended tasks gradeable via rubrics, supporting both rubric optimization and downstream objective optimization. A companion cookbook and runnable Python examples followed in the thread.
More from Research
- DNA Typewriter reconstructs mouse embryo lineage: 1.34M profiled cells from zygote to E13.5 — anshulkundaje · 2026-10-09
- Yacine calls for a code reuse benchmark to hill-climb model behavior — yacinelearning · 2026-10-09
- Data companies are becoming research labs, with verifier design the most climbable problem — madhavsinghal_ · 2026-10-09
- Amazon runs Karpathy's AutoResearch at production scale for 12 weeks, finds 5 failure modes — amazon · 2026-10-09
- ReGain: training-free fix restores subject fidelity lost from personalizing on synthetic images — UIUC-CS · 2026-10-09
- OpenAI's claimed Navier–Stokes solution reportedly doesn't match its Lean verification — kyan100 · 2026-10-09