Poolside’s laguna s 2.1 claims 118B MoE, 1M context, and 78.5% on SWE-bench Multilingual
ben_burtenshaw · x · 2026-07-22
The post says you can run poolside’s laguna s 2.1 in llama.cpp by using poolside’s fork, then launch llama-server with the laguna-s-2.1-Q4KM.gguf model.
It also shows how to enable DFlash speculative decoding with a second draft model and lists the model’s headline claims: 118B MoE, 8B active, 1M context, open weights, OpenMDW license, plus benchmark results of 70.2% on terminal-bench 2.1, 78.5% on SWE-bench Multilingual (roughly level with Claude Sonnet 5), and 59.4% on SWE-bench Pro.
The model is presented as able to run on a single DGX Spark and support up to 24 hours of autonomous runs.
Related event: Poolside Releases 118B Open-Weight Coding Model Laguna S 2.1(35 posts)→
More from coding & agent
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- Steal this idea: prompt-to-hardware where agents assemble custom devices — paraschopra · 2026-09-11
- Model Is the Least Interesting Part: A Guide to Six Core AI Architectures from RAG to Multi-Agent — goyalshaliniuk · 2026-09-11
- Non-coder builds layered memory architecture: 20k tokens tracks a year of agent conversations — matteoianni · 2026-09-11
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11