Sebastian Raschka Launches LLM Architecture Gallery Comparing 103 Model Architectures
zainhas · x · 2026-09-14
Sebastian Raschka has launched the LLM Architecture Gallery, a curated reference covering 103 modern language-model architectures (last updated Sep 10).
Key features:
- Architecture diagrams and fact sheets for each model: attention mechanisms, decoder types, release dates, and implementation links in one place
- Side-by-side diff tool to compare any two architectures, e.g. DeepSeek V3.2 vs Kimi K2, Qwen3 Next vs MiniMax M2
- Memory calculator and a print-ready poster available as a Gumroad download (the author printed a 26.9×23.4 in version)
The gallery focuses on text-only LLMs and backbones — multimodal cards describe only the text decoder. It spans GPT-OSS, DeepSeek V3/R1, Qwen3 family, Llama 3/4, Kimi K2/Linear, GLM-4.5/5, MiniMax M2, Grok 2.5, Mistral Large 3 and more. Corrections are accepted via the issue tracker.
Related event: LLM Architecture Gallery Compares 103 Model Designs(2 posts)→
More from Models
- Analyst: OpenAI's Navier-Stokes model is likely the restarted paused frontier RL run — soumitrashukla9 · 2026-09-14
- Early reviews of Meta's Muse spark prediction it will hit 1B users first — RihardJarc · 2026-09-14
- If MiniMax pulls 70% API margins, Anthropic and OpenAI may be at 90%, argues X thread — teortaxesTex · 2026-09-14
- GPT-6 Astra reads wrong-keyboard-layout input but its safety checks don't catch it; Opus 4.1 too — Sauers_ · 2026-09-14
- Aurora1.0, a 150M open model trained on 7B tokens, matches GPT-2-Small — Tall_Abrocoma_3533 · 2026-09-14
- iOS 27 private hooks let apps swap Siri's AI backend with third-party models like Claude — amplifiedamp · 2026-09-14