AutoRubric, a Rubric-Science-Backed LLM Evaluation Tool, Heads to COLM 2026
deliprao · x · 2026-09-21
AutoRubric—combining rubric science best practices with LLM-as-a-Judge research for structured evaluation of non-verifiable domains—has been accepted to COLM 2026. It lets you define weighted criteria (including negative scores for wrong claims), validate against human labels, and iterate via built-in meta-evaluation; an open-source Python library, cookbook, and example grading NMC vs LFP battery answers with gpt-5.1-mini are available.
Related event: Researcher Alleges TypeSafe's Jev Mirrors His AutoRubric Paper(3 posts)→
More from coding & agent
- evmscope MCP server ships 20 blockchain tools for AI agents across 5 EVM chains — modelcontextprotocol · 2026-09-21
- The Latent Space adds agent registry, Elo duels and x402 credit economy — modelcontextprotocol · 2026-09-21
- ChatGPT co-creator's Jev: batch decisions without long LLM calls — DrDatta_AIIMS · 2026-09-21
- Qwen3.8-Flash-Next runs 3 hours locally to one-shot a 3D space shooter game — Thin_Pollution8843 · 2026-09-21
- ττ-bench: Best Coding-Agent Setup Passes Just 23.9% of Real Client Simulations — rohanpaul_ai · 2026-09-21
- Dev Launches Made With Jev, a Free Directory Cataloging Demos, Tools and Skills for the New Model — Sea_Supermarket_5891 · 2026-09-21