Dev Dissects GLM-5.3 Flash Weights: 320B Params, 18B Active, 12,096 Expert Cells Measured
teortaxesTex · x · 2026-09-03
Developer superalesha ripped apart the deployed GLM-5.3 Flash NVFP4 checkpoint on 4x RTX PRO 6000 and visualized it in Weight Atlas — measured, not read from config: 320B total / 18B active params, 45 language layers, a 24-block vision tower, and 12,096 expert cells (42 layers x 288 experts), each with exact REAP importance from 12.59M tokens per layer, route share, and output contribution, sliceable across 14 domains including Russian and vision.
More from Models
- IBM releases Granite 4.2: free open-source models built for AI agents, runs locally — krvarshney · 2026-09-03
- Astra's rumored looped transformer gets a technical debunk, with Oriol Vinyals citing Universal Transformer — OriolVinyalsML · 2026-09-03
- Muse model now testable in opencode, Cursor support still uncertain — talkaboutdesign · 2026-09-03
- Grok 4.7 reportedly lands in 10 days: ~2.1T params, 40% scale jump, claims to top all models — tetsuoai · 2026-09-03
- Baseten ships GLM-5.3 Fast: speed-optimized open-weight model for real-time workloads — baseten · 2026-09-03
- Users Report Claude Racking Up Daily Mistakes and Hallucinations — lilyraynyc · 2026-09-03