Bad Judges Can Fool Consistency Metrics
sanmikoyejo · x · 2026-07-12
A forwarded thread mentions that a deliberately flawed judge can still achieve a human Spearman score of around 0.45, but fails the weak monotonicity test (12.5%). This indicates that simply looking at human consistency is not enough to detect certain systematic failure modes.
More from Research
- Style-similarity analysis puts Kimi K3 closer to Claude Fable 5 than to K2.6 — soumitrashukla9 · 2026-07-21
- A GLP1R variant may explain stronger Ozempic weight loss, and the team built an agent workflow — julia_kiseleva · 2026-07-21
- Proceedings for the second geometry-grounded representation learning workshop are now online — erikjbekkers · 2026-07-21
- New survey maps how agentic systems are learning to improve themselves — SchmidhuberAI · 2026-07-21
- A curated TTS list for voice agents tracks latency, cancellation, and evals — mahimairaja · 2026-07-21
- Jacob Tsimerman interview frames LLMs as a turning point for mathematical discovery — stevenstrogatz · 2026-07-21