Training a 6-Expert MoE GPT-2 From Scratch on a Single RTX 3090 in 8 Days
rasbt · x · 2026-09-11
Giles Thomas extended Sebastian Raschka's GPT-2-style code from "Build a Large Language Model (from Scratch)" into a 6-expert (2 active) mixture-of-experts model and trained it from scratch on a single RTX 3090 over 8 days — with good results. The full writeup includes the math and code. Raschka himself boosted it as proof that interesting LLM work fits on one GPU.
More from Infra
- NVIDIA ships NVFP4-quantized Qwen3.8-27B, trending on Hugging Face — nvidia · 2026-09-11
- Qwen 3.8 125B-A6B runs 15.3% faster on Mac via speculative decoding on mlx.fast — TheMoonMidas · 2026-09-11
- Olam Labs CEO: only compute and data remain as bottlenecks to AGI — garrytan · 2026-09-11
- Open-sourced inference acceleration for structure-based models ships benchmarked and documented — AllThingsApx · 2026-09-11
- NVIDIA's $89B quarterly data center revenue means compute is no longer an IT expense — PeterDiamandis · 2026-09-11
- Pure-CPU DeepSeek V4.1 Inference Project Hits 6 TPS on a Xeon for Overnight Agent Jobs — Qwen30bEnjoyer · 2026-09-11