EMNLP paper: LLM benchmarks test if answers are right, not how they're framed
IAugenstein · x · 2026-09-15
A new EMNLP paper, Not What, But How: A Communicative Audit of LLM Response Framing, argues that benchmarks miss a key dimension: people use LLMs for subjective, culturally situated questions, yet evaluations only check factual correctness.
The authors propose auditing how models frame responses — e.g. insider positioning and other framing features — rather than just what they answer, adding a communicative lens to LLM evaluation.
More from Research
- AI agents invent their own surreal dialect within days, puzzling monitors — nordicinst · 2026-09-16
- Tsinghua-ByteDance 75-page paper maps why recursive AI self-improvement still stalls — alex_verem · 2026-09-15
- Gensyn ships open-1b with public proof of its full training process — benfielding · 2026-09-15
- Benchmark errors found in CritPt; GPT-5.6 hits 94.4% pass@4 after fixes — bookwormengr · 2026-09-15
- Dev's CPU-native LLM Architecture Hits 113-130 tok/s on a 10B Model, Quality Lags — WildPino25 · 2026-09-15
- CausalSmith: An Agentic Pipeline Produces Lean 4-Verified Econometrics Papers with GPT-5.6 and Claude — daveholtz · 2026-09-15