ACL'23 outstanding paper: discriminative LMs may generalize better than autoregressive models

ysu_nlp · x · 2026-09-24

The author recaps an ACL'23 Outstanding Paper arguing Jev-like discriminative models matter for agents: language models are originally discriminative (a distribution over all sentences); autoregressive factorization forces token-level, local decisions that generalize poorly in new environments. Empirically, autoregressive ArcaneQA shows skewed distributions over seen data while discriminative Pangu aligns seen/unseen probability distributions far better.

Related event: ACL'23 Distinguished Paper: Discriminative Models May Generalize Better Than Autoregressive Ones(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →