Cognition's SWE-2 outperforms Fable-5.1 and Opus 5.5 on Animation Bench, testers report
himanshustwts · x · 2026-10-06
A tester reports that Cognition's SWE-2 beat Fable-5.1, 6-Sol, and Opus 5.5 on Animation Bench in an overnight evaluation — a surprise to the team, who shared results with Cognition on Slack. The cited tweet adds that "Devin is actually great now." Excitement is building for SWE-2.1.
More from coding & agent
- Google's AIM paper: research agents improve faster by mapping and auditing ideas, beating baselines up to 3.1x sooner — rohanpaul_ai · 2026-10-06
- Most Lawyers Use AI Legal Tools Only at Basic One-Shot Prompt Level, Says Attorney — jkubicki · 2026-10-06
- GPT-2 × Codex builds 'alien eye' tracker web app that admits when it can't see — MikePFrank · 2026-10-06
- Codex tasks widget broken? Editing it to select a specific host fixes it for now — Dimillian · 2026-10-06
- xAI TypeScript SDK hits v0.2.2: retryBeforeOutput now retries create() without stream — tetsuoai · 2026-10-06
- Hugging Face Kernels quickstart: load GPU-optimized kernels in one line — ariG23498 · 2026-10-06