Autorubric ships 25-recipe cookbook for rubric design, judge calibration and cost control
deliprao · x · 2026-10-09
The Autorubric team released a cookbook with 25 runnable recipes covering rubric design, judge calibration, ensembles, cost control, and agent skill improvement.
Key points:
- Two judge kinds: LLM judges (one criterion per call, or the whole rubric via llmcalls="peritem") and decision-model judges like TypeSafe's Jev that return probabilities in one request.
- Cascades and ensembles: mix both, or let a decision model take the first pass and hand uncertain criteria to an LLM for cheap grading.
- Tiered content: Tier 1 foundations (first eval, dataset management, explanations), Tier 2 reliability (few-shot calibration), plus validation/production and specialized scenarios.
Engineers doing evals and reward modeling can follow the recipes hands-on today.
More from coding & agent
- One Prompt, All Code: Claude Motion + Higgsfield Katana Demo Wows Users — SimplyAnnisa · 2026-10-09
- Antigen launches offensive security agent that probes, exploits and auto-patches systems — alishbaimran_ · 2026-10-09
- Yacine calls for a code reuse benchmark to hill-climb model behavior — yacinelearning · 2026-10-09
- Amazon runs Karpathy's AutoResearch at production scale for 12 weeks, finds 5 failure modes — amazon · 2026-10-09
- Neo4j's Road to NODES offers 6 free GraphRAG and agent workshops starting October 1 — JeremyCMorgan · 2026-10-09
- Engineer jokes some devs will never know the pleasure of being shredded in human code reviews — KevinNaughtonJr · 2026-10-09