Leaked Meta 30B Open Model Architecture Fuses Gemma 4 and Vision

ivan_bezdomny · x · 2026-08-10

Meta is reportedly returning to open-weight models, starting with a 30B dense model followed by muse spark 1.2.

Architecturally, the model uses knowledge distillation from muse spark and resembles Gemma 4 (Llama 3 + SWA + vision encoder), featuring scale-free QK Norm and tanh soft-capping. Compared to Gemma 4, it has fewer layers (52 vs 60), significantly more vision layers (50 vs 27), a 2x larger SWA window (2k vs 1k), and half the attention width (4k vs 8k).

Related event: Meta Returns to Open Source with 30B Parameter Model(3 posts)→

Original post →

More from Models

Models channel →