Single-metric robustness claims for LLMs can mislead, multi-level arXiv study finds

burny_tech · x · 2026-09-06

The paper 'How Perturbations Propagate' tracks six natural and synthetic input perturbations (typos, token shuffling, gradient-guided HotFlip) through decoder-only LLMs at three levels: output behavior, hidden-state geometry, and attention-head function.

Original post →

More from Research

Research channel →