Cognition launches SWE-2 claiming frontier-level coding at 70% lower cost, but skips Terminal-Bench 4
steipete · x · 2026-09-11
- Cognition officially introduced SWE-2, its model closest yet to the frontier, claiming parity with recent frontier models on leading evals at up to 70% lower cost.
- The team says it scaled RL to multiple trillions of parameters, pushing the Pareto curve on both capabilities and cost.
- The thread's real point is a callout: Cognition conveniently posted its Terminal-Bench 2.1 score while omitting Terminal-Bench 4, prompting criticism that the announcement cherry-picks favorable benchmarks.
More from Models
- Developer: Astra is a good model — jxnlco · 2026-09-11
- New Model Hits Opus-Level Benchmarks at Wild Efficiency, RL Infra Details Emerge — nrehiew_ · 2026-09-11
- V4.1's Reasoning Curve Isn't Linear; Agent Teams Explicitly RL-Trained to Collaborate — nrehiew_ · 2026-09-11
- V4.1 Is the First Model With a Numerical Reasoning Effort Parameter Tied to Length Penalty — nrehiew_ · 2026-09-11
- Only 15 Fused Kernels in Prefill: V4.1's Inference Stack and Minute-Scale SWA Cache — nrehiew_ · 2026-09-11
- DeepSeek V4.1 Flash post-training is fully data-centric: synthetic multi-agent trajectories plus checkpoint merging — nrehiew_ · 2026-09-11