300B Tokens = 10,000 Years of Nonstop Speech: Fleuret Puts Training Scale in Perspective
francoisfleuret · x · 2026-09-09
Former Meta research scientist Francois Fleuret quantifies LLM training corpora: at 8,000 words spoken per hour, 300B tokens (230B words) equals about 10,000 years of speaking 8 hours a day — a vivid illustration of how much data modern models consume.
More from Research
- ICML 2026 outstanding paper drama: concurrent diffusion sampling result, only one got the award — peter_richtarik · 2026-09-09
- UrbanLLMind: 10k memory-equipped LLM agents simulate a week of real San Francisco movement — anas_ant · 2026-09-09
- Radial Science commits $20M to Prism to make protein motion measurable and actionable — anshulkundaje · 2026-09-09
- OpenAI claims Navier-Stokes Millennium Prize solution by agent group running next-gen model — TinfoilTricorn · 2026-09-09
- The synthetic data bible: recommended starting point — yacinelearning · 2026-09-09
- 2.3M Danbooru tags corrected via human review and released as open dataset — grio43 · 2026-09-09