Single-metric robustness claims for LLMs can mislead, multi-level arXiv study finds
burny_tech · x · 2026-09-06
The paper 'How Perturbations Propagate' tracks six natural and synthetic input perturbations (typos, token shuffling, gradient-guided HotFlip) through decoder-only LLMs at three levels: output behavior, hidden-state geometry, and attention-head function.
- Evaluated across four GPT-2 and two Qwen2.5 checkpoints using CKA and intrinsic dimension, plus attention-head analysis in GPT-2.
- Perturbation types yield distinguishable metric profiles not fully captured by output measures and only partly consistent across checkpoints; copying scores correlate with activation-patching recovery under substitution/shuffling.
- HotFlip perturbations cause stronger behavioral and representational disruption than rate-matched random substitutions, consistent across all six checkpoints.
- Conclusion: robustness claims from a single metric can mislead; multi-level evaluation is needed.
More from Research
- AI could crack Navier-Stokes on its own — and add almost no value to math — NathanpmYoung · 2026-09-06
- Stanford cs336 lectures give a shoutout to NoPE research — xhluca · 2026-09-06
- Tiny 1.5B local agent stops being confidently wrong with source-tier verification, finds real bug — UzairArain554 · 2026-09-06
- MasonKamb: gradient descent is the 'original sin' behind LLM-human cognition divergences — _arohan_ · 2026-09-06
- Chris Potts' IPAM talk on interpretability and subliminal learning now available — ChrisGPotts · 2026-09-06
- Prove2Me: the Lean crowdsourcing platform behind Anthropic's Fermat's Last Theorem formalization — burny_tech · 2026-09-06