Poolside’s Laguna S 2.1 118B release includes full eval trajectories, drawing praise
_lewtun · x · 2026-07-22
A researcher praises Poolside for publishing the full trajectories behind its evaluations, saying it is valuable for understanding how the scores were obtained and, in particular, how the system prompt affected MATH results.
The quote refers to Laguna S 2.1, a new 118B-parameter model. The author says this kind of transparency should become standard practice, because reproducing other vendors’ eval scores often costs a lot of time and GPU budget.
More from Models
- Google Exec Seeks Feedback on Gemini 3.6 Flash & 3.5 Flash-Lite Performance — patloeber · 2026-07-22
- Gemini 3.6 Flash is 2x faster and 18% cheaper, but independent tests say it is not smarter — etherd0t · 2026-07-22
- Critic says OpenAI incident coverage confuses bad reward functions with autonomy — ambaonadventure · 2026-07-22
- Google Launches Gemini 3.5 Flash Cyber Model for Security Teams — pushmeet · 2026-07-22
- Rumor says GPT-5.6 Sol could hit 750 tok/s after Cerebras upgrades — haider1 · 2026-07-22
- LeCun reposts Hugging Face’s case for open-weight models in cyber defense — ylecun · 2026-07-22