From Kaggle Champion to ULMFiT: How Jeremy Howard Rewrote Language Model Training
bigaiguy · x · 2026-10-06
A long thread recounts Jeremy Howard's career, arguing that in early 2018 he showed the way people trained language models was backwards, and within months OpenAI and Google built breakthrough models on the same idea.
- Born in Australia, he studied philosophy, worked as a management consultant, and co-founded the email service Fastmail. He moved into data science, became the top-ranked Kaggle competitor and then its president and chief scientist, and founded Enlitic to apply deep learning to medical imaging.
- In 2016 he and Rachel Thomas started fast.ai, premised on doing serious deep learning without a PhD or a datacenter. Their free Practical Deep Learning for Coders course had students building real models from lesson one, and thousands used it to change careers.
- In January 2018 he and Sebastian Ruder published ULMFiT: pre-train a language model on large general text, then fine-tune it on a specific task. Transfer learning was standard in computer vision but had never worked reliably for text. ULMFiT cut error by 18–24% on most tested datasets, and with only 100 labeled examples matched models trained from scratch on 100x more data.
- OpenAI released GPT-1 in June 2018 and Google released BERT that October, both built on the pre-train-then-fine-tune recipe that became the field's foundation.
- In 2023 he co-founded Answer AI with Eric Ries; in 2024 the lab worked with Tim Dettmers to enable training very large models on consumer GPUs. He also proposed llms.txt to help websites talk to language models.
His courses are free and his library is open source. The author concludes that the most influential idea in modern language AI came from someone whose main job was teaching people how to use it.
More from Companies & People
- Zuckerberg says Llama 4 failed because it was staffed like Instagram, not a frontier lab — rohanpaul_ai · 2026-10-06
- Scoop claims Anthropic offered NDAs and pay to top probabilists, who declined — ctjlewis · 2026-10-06
- Sam Altman: 'Superintelligence' is a more accurate term than 'artificial intelligence' — 4KTV · 2026-10-06
- Chris Lattner rebuilt the compiler stack for 25 years: LLVM, Clang, Swift, MLIR, Mojo — thisdudelikesAI · 2026-10-06
- OpenAI 'acting like cornered animals' amid usage-limit backlash; exec declines comment citing safety — CtrlAltDwayne · 2026-10-06
- HKMA grills HSBC on why its new AI hub is in Singapore, not Hong Kong — AIFlow_ML · 2026-10-06