Minimind: Train a 64M-parameter LLM from scratch in just 2h
jingyaogong · github · 2026-08-31
GitHub project jingyaogong/minimind provides a complete codebase and tutorial to train a small 64M-parameter Large Language Model from scratch in just 2 hours. It is suitable for learning LLM training workflows and principles.
More from Research
- AlphaEvolve sets new record for matrix multiplication exponent: omega to 2.371177 — 新智元 · 2026-09-01
- Building a long-term memory benchmark for agents: what to add? — True_Mongoose_7073 · 2026-09-01
- Building a high-recall, traceable "second brain" RAG system? — iMiguelmars · 2026-09-01
- Study reveals cross-layer activation patterns in hybrid attention models — 机器之心 · 2026-09-01
- Group-averaged Markov chains papers updated, blending group theory with Markov chains — michaelchchoi · 2026-09-01
- The agent doom loop isn't the model being dumb — it's the transcript working against you — RunAI_Coder · 2026-09-01