Adding one JSON field degraded LLM judge stability: 18/21 runs shifted, reordering didn't help
alexpran · reddit · 2026-09-23
- An open-source eval tool maintainer reports his single-score judging prompt agreed with his labels on 16/21 cases in 18 of 21 runs — as stable as these things get.
- Adding a second field to the same JSON reply didn't collapse agreement but loosened it: median intra-sample disagreement jumped from 4 to 7-8 in every run, with zero overlap with the old range.
- He hypothesized generation order (model invents a description before scoring) and moved the field after the score — numbers were identical. Rerunning the old prompt the same day gave 6/21, inside its usual band, ruling out model drift.
- No mechanism offered, only the effect. He's asking whether anyone has measured how existing output changes when new fields are added — a useful empirical data point for agent eval engineering.
More from coding & agent
- Codebase-Memory MCP server hits 44k stars: 158 languages, 99% fewer tokens — DeusData · 2026-09-23
- Strands open-sources Harness SDK for building production AI agents in Python and TypeScript — strands-agents · 2026-09-23
- Entire Video Edit Costs $0.03 With No Editor, No Claude Code, No Paid APIs — CurieuxExplorer · 2026-09-23
- Studio Jadu tests its AI assistant with annoying, impatient synthetic users — andimarafioti · 2026-09-23
- Fine-tuning a 194M GLiNER2 tag classifier on HF Jobs for ~$1.50 lifts accuracy from 10% to 69% — vanstriendaniel · 2026-09-23
- Dev builds 3D Spider-Man game from scratch with Opus 5.5 in Three.js — TawohAwa · 2026-09-23