LLM-as-judge without calibration is just opinion laundering
bgoncalves · x · 2026-08-28
Bruno Gonçalves argues that using LLM-as-judge without calibrating against human labels isn't evaluation; it's "opinion laundering."
- Core Argument: It's just replacing one shrug (uncertainty) with a more expensive one.
- Implication: Model scores lacking Ground Truth calibration may merely mask evaluation uncertainty.
More from Research
- Handbook on RAG & Context Engineering Based on Real Papers — techNmak · 2026-08-28
- Developer Publishes Handbook on RAG and Context Engineering Based on Real Papers — techNmak · 2026-08-28
- Schmidhuber: 1st backprop-trained CNN for vision from 1988 — SchmidhuberAI · 2026-08-28
- Implementing a modern LLM runtime in 700 lines of C — Critical_Physics8 · 2026-08-28
- Qwen3-8B fine-tuned with LoRA+GRPO mimics a famous ML blogger, fooling every AI detector — OtherRaisin3426 · 2026-08-28
- Fieldwork data is messy: AI challenge of unstructured reconstruction — anthara_ai · 2026-08-28