Technion: generalization is stability, not accuracy — cross-dataset variance can reverse LLM rankings

Technion · hf · 2026-10-02

Technion researchers argue LLM generalization should be measured as output consistency across paraphrases of the same input, not aggregate accuracy on a single prompt format. They introduce SAGO (Stability-Aware Generalization Objective), which quantifies behavior variability across generation consistency, internal activations, confidence, and response mirroring. Findings: commonly used models exhibit statistically significant generalization instability, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.

Original post →

More from Research

Research channel →