Is DeepSWE dead? Coding benchmark changelog untouched since September 3
rgb328 · reddit · 2026-10-10
A Reddit user notes the DeepSWE coding benchmark appears stalled: recent models including Claude 5.5, sol 6.1, and Mistral Large 4 have no results, and the changelog at deepswe.datacurve.ai hasn't been updated since September 3. No official word on whether the effort is discontinued.
More from Models
- Anthropic starts frequent behavior reports, detailing four unintended Claude actions — AnthropicAI · 2026-10-10
- Cloudflare releases clef-omni, an open omni-modal model with audio, image and video input — ritakozlov · 2026-10-10
- Strong backbones plus light fine-tuning beat synthetic data, says researcher whose model tops benchmarks — antoine_chaffin · 2026-10-10
- AWS Bedrock posts legacy notices for Claude Opus 4.1, Sonnet 4 and Sonnet 4.5 — repligate · 2026-10-10
- Gemini web shows low/medium/high thinking effort tiers, hinting Gemini 4 Argon is near — teocutter · 2026-10-10
- Google slammed for not releasing Argon after officially announcing it — almmaasoglu · 2026-10-10