Technion: generalization is stability, not accuracy — cross-dataset variance can reverse LLM rankings
Technion · hf · 2026-10-02
Technion researchers argue LLM generalization should be measured as output consistency across paraphrases of the same input, not aggregate accuracy on a single prompt format. They introduce SAGO (Stability-Aware Generalization Objective), which quantifies behavior variability across generation consistency, internal activations, confidence, and response mirroring. Findings: commonly used models exhibit statistically significant generalization instability, behavioral axes capture independent failure modes, and cross-dataset variation can reverse model rankings.
More from Research
- LOCI: hybrid spatial linear memory lets streaming world models recall revisited scenes at ~30% less memory — IFM · 2026-10-02
- SAKIKO auditing shows +55 net-gain interventions corrupt over half of correct tool-using LLM decisions — UniversityofBirmingham · 2026-10-02
- Researchers warn AI-written 'salami' papers are flooding arXiv with low-value work — EhudReiter · 2026-10-02
- Arena Physica explains why FEM solvers never compute E-fields at mesh nodes — burny_tech · 2026-10-02
- AWSM grounds LLM-agent 3D scene reconstruction in IMU, depth and pose evidence — anselm · 2026-10-02
- Study links attention nonlinearity to power-law massive activations and scaling laws — burny_tech · 2026-10-02