HCOMP best paper: ROUGE and LLM judges fail to measure summary reader satisfaction
mdredze · x · 2026-10-09
Isabel Cachola et al.'s "Information Satisfaction" won HCOMP Best Paper, exposing a core gap in summarization evaluation: a summary should serve a specific reader, yet mainstream metrics can't measure that.
- Proposes an "information satisfaction" axis: the same topic serves a biomedical researcher and a family doctor differently; short queries can't capture this, while stable reader personas are a practical signal
- Most popular metrics, including strong LLM-as-judge ones, fail basic perturbation tests and don't track reader satisfaction
- Validated with expert human evaluation
Paper: arXiv 2608.14457
More from Research
- Inherit-MAS cuts multi-agent token use by up to 34.6% with evolution-inspired inheritance — Songtao Wei · 2026-10-09
- Zero human labels: auto-generated soccer tracking dataset hits 62.5 HOTA, beating fine-tuned FairMOT — RexDouglass · 2026-10-09
- Why AI doesn't actually read words: from BPE subwords to byte-level models like BLT and H-Net — jbhuang0604 · 2026-10-09
- NVIDIA details HSTU recommender inference stack with up to 5.93x lower latency — PyTorch · 2026-10-09
- Preprint: LLMs store numbers as curves and helices, but compute comparisons differently — tweetsatpreet · 2026-10-09
- Best model was cheapest: open-weights model ran 669 clinical decisions for 1.7 cents — antoine_chaffin · 2026-10-09