Raschka: with $100M to build an LLM, I'd skip pretraining and invest in post-training
rasbt · x · 2026-09-20
- Sebastian Raschka shares that if given $100M to build a state-of-the-art LLM, he would use an existing model and put the budget into post-training.
- The thread also mentions Jev's impressive generalization, reportedly trained with 100% synthetic data and no pretraining—only post-training on an existing model—saving substantial costs.
Related event: Raschka would spend a $100M LLM budget entirely on post-training(3 posts)→
More from Models
- Encoder-style classification gets hot again: one multimodal BERT solved 1,000+ classification tasks — cwolferesearch · 2026-09-21
- Deployers told: benchmarks are a directional signal, not a substitute for your use case — evijit · 2026-09-21
- Vehicles for Claude tiers: Haiku is a bicycle, Opus a C130, Fable the mothership — __rum_ham__ · 2026-09-21
- Unverified rumor: OpenAI's mysterious 'Bell' model targets narrow ASI in math and code — VraserX · 2026-09-20
- Xiaomi's MiMo V2.6 RL training livestream burns $3.24M in 4.5 days — teortaxesTex · 2026-09-20
- Resetwatch plugin aggregates usage limits and reset times for 14 AI providers in one page — Teknium · 2026-09-20