Judge simulation study is now testing whether kappa-style agreement changes trust
IanArawjo · x · 2026-07-21
The author says they have **not yet tried binary data** for the judge simulations. - They are first expanding the simulations and breaking results down by **agreement rate**. - The goal is to test whether reporting agreement such as **kappa = 0.85** actually changes how much trust people place in the judge. - The exchange suggests the underlying project is about **evaluating judge reliability**, not just producing a single score.
Related event: Study Warns of High False Positive Rates in Unchecked LLM Judges(7 posts)→
More from Research
- Gemma 4 12B visualization shows what the model predicts from video patches — arjunrajlab · 2026-07-21
- Baseten study finds new facts in LLM weights are fragile unless trained from many restatements — alex_verem · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21
- uv-scripts/ocr returns to the top of Hugging Face datasets with a JSON model picker — vanstriendaniel · 2026-07-21
- DeepSearch-World trains web agents with 420K verifiable QA tasks — HKUST · 2026-07-21
- GigaAM Multilingual targets low-resource Central Asian ASR with 2M hours of audio — ai-sage · 2026-07-21