Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve
YouJiacheng · x · 2026-09-11
- Cognition explains that SWE-2 is its first model to support effort levels, enabled by significant improvements to the length penalty recipe.
- In a single RL run, the team pushes the Pareto curve while preserving its shape: medium effort becomes both cheaper and smarter, while max effort learns to use more tokens and turns to achieve the highest scores.
- @RichardYRLi highlights the novel use of a KKT (duality) argument in modern LLM RL applications.
- More details in Cognition's blog post.
More from coding & agent
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- Steal this idea: prompt-to-hardware where agents assemble custom devices — paraschopra · 2026-09-11
- Model Is the Least Interesting Part: A Guide to Six Core AI Architectures from RAG to Multi-Agent — goyalshaliniuk · 2026-09-11
- Non-coder builds layered memory architecture: 20k tokens tracks a year of agent conversations — matteoianni · 2026-09-11
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11
- 9-year backend dev: AI code isn't the problem, the rate of making a mess is — Sweaty-Landscape-561 · 2026-09-11