Training a 6-Expert MoE GPT-2 From Scratch on a Single RTX 3090 in 8 Days

rasbt · x · 2026-09-11

Giles Thomas extended Sebastian Raschka's GPT-2-style code from "Build a Large Language Model (from Scratch)" into a 6-expert (2 active) mixture-of-experts model and trained it from scratch on a single RTX 3090 over 8 days — with good results. The full writeup includes the math and code. Raschka himself boosted it as proof that interesting LLM work fits on one GPU.

Original post →

More from Infra

Infra channel →