Experts Rise Where LLMs Disagree: rationale labeling cuts codebook revision from months to days
windx0303 · x · 2026-09-26
Cornell researchers (Zeyu He, Ting-Hao 'Kenneth' Huang, et al.) propose using cross-LLM disagreement to target expert attention when revising annotation codebooks for large-scale text labeling.
Method
- Apply an early codebook with LLM annotators, surface cases with strong cross-model disagreement, and elicit expert feedback.
- Three feedback modes compared: (i) Codebook Verifying (editing LLM-generated revisions), (ii) Question Answering about disagreements, (iii) Rationale Labeling of disagreement cases.
Results (thousands of tutoring-session transcripts)
- Rationale Labeling performed best: 64.9% LLM-labeling accuracy vs expert labels, beating the expert-revised codebook (57.8%).
- Best Question Answering setting also outperformed it at 60.5%.
Takeaway: LLMs can strategically direct expert effort, shrinking months of codebook revision to days without sacrificing labeling quality.
More from Research
- Stanford's Noah Goodman uses philosophy to improve LLM pretraining, jokes ASI achieved — xuanalogue · 2026-09-26
- Researchers surface spurious probes across models: Sonnet 5 recommends green tea in evals, oolong in production — jankulveit · 2026-09-26
- Three papers, one warning: 1% synthetic data can trigger strong model collapse — suchenzang · 2026-09-26
- Simulation beats distillation: the real story of synthetic data is post-training worlds — realsohamparekh · 2026-09-26
- Dev hails continual learning paper: AGI defined in 2000, only now is anyone training for it — willcb · 2026-09-26
- Lawrence Krauss podcast asks whether AI will supercharge scientific paper mills — willcb · 2026-09-26