TI_var Predicts Judge Alignment
sanmikoyejo · x · 2026-07-12
The thread shares TIvar's predictive results for judge-level human alignment: on the validation set, the Spearman correlation is 0.833 for average labels, 0.905 for pass1, and 0.738 for pass2. Experiments cover judge families including GPT-5.3-Codex, Claude Opus 4.7, Qwen3-Coder.
More from Research
- Fast ViT shows strong ImageNet results; scaling runs needed next — ducha_aiki · 2026-09-11
- Loss Functions Are Scientific Assumptions: MSE Implies Gaussian Noise, Cross-Entropy Implies Bernoulli — bravo_abad · 2026-09-11
- SymKit MCP: 44 tools for AI agents to verify symbolic derivations — Foreign-Specific-604 · 2026-09-11
- Researchers: LLMs under pressure invent new languages unreadable to humans — mikeflache · 2026-09-11
- Mi-Ripple fixes ripple artifacts left by iterative AI image editing — Miyang-AI · 2026-09-11
- DRG-MAPPO uses dynamic role graphs to boost multi-agent air combat win rates — China666 · 2026-09-11