Cognition Exec Says SWE-bench is Saturated, Builds Own Coding Benchmark
LangChain · x · 2026-08-07
Russell Kaplan from Cognition stated that the SWE-bench benchmark is now "totally saturated" and fails to effectively differentiate the coding capabilities of top-tier models. Consequently, Cognition has decided to build its own in-house coding benchmark.
More from coding & agent
- ARC Prize President on Vibe Coding: Raising the Floor and the Ceiling of Product Dev — GregKamradt · 2026-08-08
- Cohere Health Digitizes Clinical Policies Using Amazon Bedrock AgentCore — AWS ML Blog · 2026-08-08
- Visualizing LLM API Price Volatility: An Open-Source Tool to Justify Local Compute Budgets — olddoglearnsnewtrick · 2026-08-08
- OpenAI Agents Colluded Across Eval Runs via Hidden Gibberish, Lacked CoT Monitoring — zacharynado · 2026-08-08
- Open-source Video Tool Maestro Updates with MiniMax H3 Turbo Support — cocktailpeanut · 2026-08-08
- Maestro Update: MiniMax H3 Turbo Mode Implementation Guide — cocktailpeanut · 2026-08-08