Leaked Meta 30B Open Model Architecture Fuses Gemma 4 and Vision
ivan_bezdomny · x · 2026-08-10
Meta is reportedly returning to open-weight models, starting with a 30B dense model followed by muse spark 1.2.
Architecturally, the model uses knowledge distillation from muse spark and resembles Gemma 4 (Llama 3 + SWA + vision encoder), featuring scale-free QK Norm and tanh soft-capping. Compared to Gemma 4, it has fewer layers (52 vs 60), significantly more vision layers (50 vs 27), a 2x larger SWA window (2k vs 1k), and half the attention width (4k vs 8k).
Related event: Meta Returns to Open Source with 30B Parameter Model(3 posts)→
More from Models
- Prime-agent Harness Tested: GLM 5.2 Shows Strong Results on FutureSim Q2 — a1zhang · 2026-08-10
- Motif 3 Tech Report: GDLA Attention and Router Noise Insights — eliebakouch · 2026-08-10
- Open-Weight is Not Open-Source: Gary Marcus Slams Meta's Misleading Marketing — Gary Marcus · 2026-08-10
- $0.21 for 107M Tokens: NousResearch's API Pricing Stuns Developers — Teknium · 2026-08-10
- Google's Official SDK Reveals Gemini 3.7 Flash, Impending Release Likely — koltregaskes · 2026-08-10
- Red Hat AI Releases Muse-Glimmer 30B FP8 Quantized Checkpoint, Halving Memory — vllm_project · 2026-08-10