A 3-month curated paper list for learning distributed LLM training and inference
East-Muffin-6472 · reddit · 2026-09-21
The author curates a beginner-friendly reading list—built over three months—covering the distributed-systems fundamentals behind LLM training and inference: distributed/tensor/pipeline/model parallelism, shared via an alphaxiv folder. They also open-sourced basic implementations of several of these methods in the GitHub repo smolcluster as reference, recommending read-code-play as the path in.
More from Infra
- Mozilla AI runs a local 30B model end-to-end to open a real bugfix PR, fully offline — mozilla-ai · 2026-09-21
- Cohere Labs launches Local AI community program for local inference and hardware tuning — Cohere_Labs · 2026-09-21
- Gewell: custom Gemma 4 inference engine cuts KV cache VRAM to 0.625x, losslessly — stoppableDissolution · 2026-09-21
- DeepSeek-V4.1-Flash redesigns the Transformer for agents, cutting KV cache to 890 bytes/token — AndLukyane · 2026-09-21
- Jev Engineering gives agents a decision brain, 193x faster and 444x cheaper in tests — agihouse_org · 2026-09-21
- UK's £225m Isambard-AI supercomputer cost about the same as one road bridge — charlieharris01 · 2026-09-21