Poolside Releases 118B Open-Weight Coding Model Laguna S 2.1

AI coding company Poolside has released Laguna S 2.1, its latest open-weight coding model. Designed specifically for agentic coding and long-horizon tasks, the model quickly topped the Hugging Face trending charts thanks to its ultra-long context support and strong benchmark performance.

Confirmed

Laguna S 2.1 utilizes a Mixture-of-Experts (MoE) architecture with 118B total parameters and 8B active parameters per token. It supports a context window of up to 1 million tokens and offers both thinking and no-thinking modes. Released under the OpenMDW license, it is built on the Transformers architecture and supports vLLM deployment. Poolside emphasized the model's stable performance in long-chain tasks and its ability to outperform some much larger models in agentic coding. Officially, it scored 70.2 on the Terminal-Bench 2.1 benchmark, rivaling or surpassing models 5 to 25 times its size. Thanks to its low active parameters, the model runs efficiently on a single NVIDIA DGX device and has day-0 support on SGLang. According to @benburtenshaw, it can already run locally in llama.cpp via a fork, allowing users to directly load quantized files. Regarding the development timeline, @Madisonkanna noted it took only 8 weeks, while @nathanbenaich mentioned a 9-week period.

Why it matters

The release of Laguna S 2.1 is more than just a new model launch; it sparks a deeper discussion about the "US open-source AI approach" and the economics of open weights. In a related interview, @Madisonkanna pointed out that this touches upon the larger question of "who is qualified to build intelligence." Its exceptional local inference efficiency also provides significant convenience for developers in actual deployment and testing.

2026-07-22 ~ 2026-07-23 · 35 related posts

Full story(2 episodes)→

Primary sources

5 near-duplicate retellings: Madisonkanna · baseten · Madisonkanna · ivan_bezdomny · nathanbenaich