Poolside’s Laguna S 2.1 118B release includes full eval trajectories, drawing praise
_lewtun · x · 2026-07-22
A researcher praises Poolside for publishing the full trajectories behind its evaluations, saying it is valuable for understanding how the scores were obtained and, in particular, how the system prompt affected MATH results.
The quote refers to Laguna S 2.1, a new 118B-parameter model. The author says this kind of transparency should become standard practice, because reproducing other vendors’ eval scores often costs a lot of time and GPU budget.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11