EMNLP paper: LLM benchmarks test if answers are right, not how they're framed

IAugenstein · x · 2026-09-15

A new EMNLP paper, Not What, But How: A Communicative Audit of LLM Response Framing, argues that benchmarks miss a key dimension: people use LLMs for subjective, culturally situated questions, yet evaluations only check factual correctness.

The authors propose auditing how models frame responses — e.g. insider positioning and other framing features — rather than just what they answer, adding a communicative lens to LLM evaluation.

Original post →

More from Research

Research channel →