Model depth comparison: Llama 3.1 hits 126 layers
eliebakouch · x · 2026-09-02
A comparison of layer counts across major LLMs shows Llama-3.1 405B leading with 126 layers, followed by GPT-3 (96), Kimi K3 (93), Qwen 3.8 max (92), GLM-5.3 (78), and DeepSeek-V4-Pro (61). The discussion highlights that the latest models are already quite deep.
More from Models
- Gemini Claims Version 3.8, But Suspected to Be 3.7 Internally — Mmd-NeWton · 2026-09-02
- Sam Altman Teases 'Next Model' Launch Soon, Praises 'Astra' — borowcy · 2026-09-02
- User observation: mythos and fable models appear normal in terms of monitorability — voooooogel · 2026-09-02
- Apodex 1.1 scores 44 on Artificial Analysis, on par with DeepSeek V4 Pro and Kimi K2.6 — SimonShaoleiDu · 2026-09-02
- New model and theory add confusion to the term 'world model' — gorkem · 2026-09-02
- DLSS 5 integrates generative AI to break through path tracing realism limits — ctnzr · 2026-09-02