Predicting Affected Questions Proves Difficult

sanmikoyejo · x · 2026-07-16

The author points out that which questions are affected by context varies drastically across different models.

They calculated the Pearson correlation for "per-sample performance changes" between models, yielding a result near zero. This indicates that this failure mode is extremely difficult to predict based solely on the questions themselves.

Original post →

More from Research

Research channel →