DeepSWE-mini: a 16-instance subset that replicates the full DeepSWE leaderboard rankings
asankhs · reddit · 2026-09-18
A new dataset, DeepSWE-mini (on Hugging Face: LocalLLaMA/deepswe-mini), is a 16-instance subset of the DeepSWE coding benchmark. Analysis showed it faithfully replicates the relative rankings (not absolute scores) of the full benchmark, letting developers quickly benchmark new local models without the full, slow run.
Related event: DeepSWE-mini Released: 16 Instances Reproduce Full Leaderboard(2 posts)→
More from Research
- Epoch AI audits 15 AI benchmarks: only 4 safe to trust at face value — Jsevillamol · 2026-09-18
- Zoom study of 176 settings: context management saves tight-window coding agents, planning just cuts cost for strong models — ZoomCommunications · 2026-09-18
- VA-Bench: Best MLLM Scores Only 53.9% Task Success on Embodied Spatial Intelligence — dalian-university-of-technology · 2026-09-18
- NVIDIA's SoL-Pi Cuts Coding Agent Token Traffic ~49% and API Cost by a Third — nvidia · 2026-09-18
- RiskChainBench: Web Agents Hit 31.9% Execution Failures in Obfuscated Abuse Chain Investigations — ZhuoXin Liu · 2026-09-18
- Vision-RL2: Region-Level RL Matches Full-Resolution MLLM Accuracy With 4x Fewer Visual Tokens — Yuheng Shi · 2026-09-18