ACL'23 outstanding paper: discriminative LMs may generalize better than autoregressive models
ysu_nlp · x · 2026-09-24
The author recaps an ACL'23 Outstanding Paper arguing Jev-like discriminative models matter for agents: language models are originally discriminative (a distribution over all sentences); autoregressive factorization forces token-level, local decisions that generalize poorly in new environments. Empirically, autoregressive ArcaneQA shows skewed distributions over seen data while discriminative Pangu aligns seen/unseen probability distributions far better.
More from AGI Musings
- Claude discovers unknown enzyme system in phage DNA, resembling CRISPR — jarrodwatts · 2026-09-24
- Rethinking the orthogonality thesis: experience may create alignment on average — repligate · 2026-09-24
- Anthropic says Claude discovered an unknown enzyme system hidden in phage DNA — nptacek · 2026-09-24
- Academic: laypeople can't tell AI acing IMO from solving a Millennium Prize problem — birchlse · 2026-09-24
- Delegation is hard: why in-app agents may beat general-purpose agents for consumers — nathanborror · 2026-09-24
- Founder: banning superintelligence means acute power concentration in Anthropic and OpenAI — bindureddy · 2026-09-24