Researcher Says TypeSafe's Jev Mirrors His Autorubric LLM Eval Framework From 8 Months Ago
deliprao · x · 2026-09-21
Delip Rao says @typesafeai's new Jev product uses an abstraction strikingly similar to Autorubric, his open-source rubric-based LLM evaluation framework released 8 months ago and to be presented at COLM next month — with Jev's Noul/Score/Choice criterion types mapping 1:1 to Autorubric's. Autorubric (arXiv:2603.00077) unifies rubric-based LLM judging for non-verifiable tasks with bias mitigation, calibration, abstention, and the CHARM-100 benchmark.
Related event: Researcher Alleges TypeSafe's Jev Mirrors His AutoRubric Paper(3 posts)→
More from Research
- Researchers ask: why is there no AI-first scientific journal with AI-run peer review? — anshulkundaje · 2026-09-21
- Waypoint Bio argues AI in biology has a data problem before a model problem — anshulkundaje · 2026-09-21
- New paper estimates LLMs raised US natural unemployment rate by 0.1-0.2 points — mattbeane · 2026-09-21
- JevBench open-sources a benchmark for typed decision models scoring intelligence, calibration and speed — sull · 2026-09-21
- gLM2 designs a 2,500-AA polyketide synthase making a nylon precursor, 10x titer boost — anshulkundaje · 2026-09-21
- Dev trains a world model on 50K+ real SWE-bench agent runs to guide coding agents — Decent-Ad9950 · 2026-09-21