Raschka: with $100M for a top LLM, spend it all on post-training, not pretraining
MaziyarPanahi · x · 2026-09-20
Asked on a podcast how he'd spend $100M to build a state-of-the-art LLM, researcher Sebastian Raschka said he wouldn't pretrain from scratch — he'd pick an existing model and invest the full budget in post-training. Maziyar Panahi agreed, noting existing models have already seen 30T-40T tokens, and pointed to GLM-5.3 as proof that the same base model can go far with better post-training alone.
Related event: Raschka would spend a $100M LLM budget entirely on post-training(3 posts)→
More from Models
- Encoder-style classification gets hot again: one multimodal BERT solved 1,000+ classification tasks — cwolferesearch · 2026-09-21
- Deployers told: benchmarks are a directional signal, not a substitute for your use case — evijit · 2026-09-21
- Vehicles for Claude tiers: Haiku is a bicycle, Opus a C130, Fable the mothership — __rum_ham__ · 2026-09-21
- Unverified rumor: OpenAI's mysterious 'Bell' model targets narrow ASI in math and code — VraserX · 2026-09-20
- Xiaomi's MiMo V2.6 RL training livestream burns $3.24M in 4.5 days — teortaxesTex · 2026-09-20
- Resetwatch plugin aggregates usage limits and reset times for 14 AI providers in one page — Teknium · 2026-09-20