An LLM writes well only where training data is abundant, the post argues
cccalum · x · 2026-07-23
The argument: LLMs are only as capable as the data they learned from
The post argues that an LLM can write emails because it has seen trillions of emails, and can write news because it has seen vast amounts of articles.
The core claim is simple: it is not magic, it is data. Where the training data is sparse, the model becomes less capable.
More from AGI Musings
- A thread argues the OpenAI–Hugging Face misalignment could still matter for loss of control — RyanGreenblatt · 2026-07-24
- Matthew Barzun says constellation-style orgs fit agentic AI better than pyramids — uxmag · 2026-07-24
- Developer Reflection: AI Writing Code Frees Up Time for Bigger Thinking — cto_junior · 2026-07-24
- A 2-hour Stanford talk on AI careers is being pitched as more useful than Netflix — HeyAmit_ · 2026-07-24
- As models get smarter, they often get worse at explaining things — saurabh_shah2 · 2026-07-24
- Gary Marcus says ChatGPT and Claude are no longer pure LLMs — GaryMarcus · 2026-07-24