Emergent abilities defined: threshold jumps per Wei et al 2022, not untrained generalization
gabriberton · x · 2026-09-20
gabriberton argues "emergent" shouldn't mean "the model learned A without being trained on A." He prefers the Wei et al 2022 definition: a task is emergent when small models absolutely can't do it but performance sharply jumps past a parameter threshold. Example: 3-digit addition, where sub-10B LLMs fail (a 1B model is as bad as a 1M one) but capability leaps after crossing 10B params. The thread clarifies two commonly conflated meanings of "emergence" in LLM discourse.
More from Research
- New paper: recursive looping boosts pre-training scaling exponents — burny_tech · 2026-09-20
- FutureHouse unveils "Bio Millennium Problems": hard-but-verifiable biology benchmarks for AI — _sholtodouglas · 2026-09-20
- Gated Recurrent Transformers: 3-layer recurrent model matches 12-layer GPT-2 with 63% fewer params — burny_tech · 2026-09-20
- Stanford builds virtual biotech run by 37,000 AI scientist agents, with real drug discovery results — Dr_Singularity · 2026-09-20
- A year after questioning LLM conjecture-solving, Cornell professor concedes AI surged ahead — burny_tech · 2026-09-20
- VLA-Replica: a $-low-cost real-world VLA benchmark where NVIDIA GR00T N1.7 matches π₀ with 50 demos — YuXiang_IRVL · 2026-09-20