Irrelevant Context Alters Model Answers

sanmikoyejo · x · 2026-07-16

This study tests how models perform when task-irrelevant context is prepended to questions across multiple QA benchmarks.

Key findings:

This highlights the models' sensitivity to contextual noise, suggesting that we must look at distribution tails rather than just average scores.

Original post →

More from Research

Research channel →