Hamel Husain: same model as judge is usually fine — verify human-label alignment first
HamelHusain · x · 2026-10-08
- Hamel Husain and Shreya Shankar answer in their AI Evals FAQ: using the same model for your main task and as LLM-as-Judge is usually fine, since the judge performs a different, scoped task.
- What matters is alignment with human judgments: aim for high TPR/TNR on a held-out labeled test set with binary classification judges, iterating against human labels.
- Only switch models if alignment fails; start with the most capable model, optimize for cost later.
More from coding & agent
- Agent harnesses are crutches; least-privilege action authorization is what matters — andreisavu · 2026-10-09
- Workers Refuse to Write Markdown Files for Agents, Seeing Knowledge Extraction — mattbeane · 2026-10-08
- Fallout: New York runs in your browser, built with Claude Opus 5.5, zero texture or sound files — chrisfirst · 2026-10-08
- When is AI automation worth it? Only for five-minute tasks that keep repeating — gethackteam · 2026-10-08
- Dad dispatches an agent to build an ESP32 music player to quit Spotify for good — natesiggard · 2026-10-08
- Andy Pavlo: AI agents create 80% of new databases and keep deleting production ones — mattturck · 2026-10-08