This week’s open-weight round: 1M-token Laguna S 2.1, 314B Motif-3-Beta, and more
rasbt · x · 2026-07-26
A roundup of several notable open-weight model releases this week, with architecture notes and early performance observations:
- Nanbeige 4.2 3B uses looped depth sharing: the same 22-layer stack is run twice, effectively doubling compute without duplicating weights. The author says the paper suggests two passes were the best trade-off, retaining about 75% token efficiency versus a standard architecture.
- Laguna S 2.1 is a 118B sparse MoE with 8B active parameters and a 1M-token context window. It is described as unusually practical on the author’s DGX Spark, using under 80 GB of RAM, though independent benchmarks are still pending.
- Motif-3-Beta is a 314B-A13B sparse MoE inspired by DeepSeek V4, but with a new attention component called Grouped Differential Latent Attention.
- Solar Open 2 is a 250B-A15B hybrid MoE that interleaves three Kimi Delta Attention layers with one GQA layer.
- Antares 1B and a smaller 0.3B variant target terminal cybersecurity tasks, built on IBM Granite 4.0 1B and trained with SFT plus GRPO.
The attached chart also compares prefill speed, decode speed, and GPU memory use against Qwen3.6 35B-A3B Q4.
More from Models
- Claude keeps defending Opus 5 instead of writing a critical review — sethlazar · 2026-07-26
- Unsloth’s Qwen3.6-35B-A3B GGUF is trending on Hugging Face — unsloth · 2026-07-26
- SOOFI may still only be tying Nemotron 3 Nano despite 2T extra tokens — JJitsev · 2026-07-26
- A simple “ask clarifying questions first” prompt gets Claude to surface missing assumptions — Commercial-Most3081 · 2026-07-26
- User says Claude Opus 5 feels more forgiving and “common-sense” than GPT-5.6 Sol — dejavucoder · 2026-07-26
- Frontier models are hallucinating less about biology, but using terms more loosely — owl_posting · 2026-07-26