Irrelevant Context Alters Model Answers
sanmikoyejo · x · 2026-07-16
This study tests how models perform when task-irrelevant context is prepended to questions across multiple QA benchmarks.
Key findings:
- Even semantically meaningless pseudo-words can cause models to significantly change their answers on certain samples
- This effect leads to both "worse" and "better" outcomes for different samples
- The upper and lower tail samples (those with the most significant changes) account for the majority of the overall variance
This highlights the models' sensitivity to contextual noise, suggesting that we must look at distribution tails rather than just average scores.
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11