Poolside’s laguna s 2.1 claims 118B MoE, 1M context, and 78.5% on SWE-bench Multilingual

ben_burtenshaw · x · 2026-07-22

The post says you can run poolside’s laguna s 2.1 in llama.cpp by using poolside’s fork, then launch llama-server with the laguna-s-2.1-Q4KM.gguf model.

It also shows how to enable DFlash speculative decoding with a second draft model and lists the model’s headline claims: 118B MoE, 8B active, 1M context, open weights, OpenMDW license, plus benchmark results of 70.2% on terminal-bench 2.1, 78.5% on SWE-bench Multilingual (roughly level with Claude Sonnet 5), and 59.4% on SWE-bench Pro.

The model is presented as able to run on a single DGX Spark and support up to 24 hours of autonomous runs.

Related event: Poolside Releases Laguna S 2.1: An 118B Open-Weight Coding Model(9 posts)→

Original post →

More from coding & agent

coding & agent channel →