The problem with LLMs as judges: Jev bets on structured math over wordy rationalizations
forevergeeks · reddit · 2026-09-30
The author argues that using LLMs as judges is inefficient: LLMs rationalize everything and are eager to please, sometimes making things up to meet expected standards.
Jev's approach to structured data stands out — still probabilistic, but returning a mathematical result beats a "diarrhea of words" explaining a decision. The post closes asking whether open-source alternatives to Jev will emerge.
More from coding & agent
- One prompt to architect a full Jev decision layer with Claude, author claims — iamrobotbear · 2026-09-30
- ChatGPT subscription now works inside 60+ partner products including Devin, OpenCode and Lovable — TheMoonMidas · 2026-09-30
- OpenAI DevDay gives AI agents their own computer — from chatting to assigning work — 141_1337 · 2026-09-30
- Agent wired to macro datasets from the 1930s watches the economy 24/7 — virattt · 2026-09-30
- Making games with Opus 5.5? You need to be Pinterestmaxxing first — nptacek · 2026-09-30
- Conductor adds Sign in with ChatGPT to bring Codex subscriptions over — charlieholtz · 2026-09-30