DeepSWE-mini: A 16-Instance Subset That Replicates the DeepSWE Leaderboard Rankings
asankhs · reddit · 2026-09-18
Redditor asankhs released DeepSWE-mini, a 16-instance subset of the DeepSWE benchmark, on Hugging Face.
- The full DeepSWE benchmark takes a long time to run, so the author analyzed it and selected a subset that faithfully replicates models' relative rankings (though not absolute scores).
- It's meant for quickly benchmarking new local models without running the full suite.
- Dataset: huggingface.co/datasets/LocalLLaMA/deepswe-mini
Related event: DeepSWE-mini Released: 16 Instances Reproduce Full Leaderboard(2 posts)→
More from Research
- Turing Post Publishes RSI 2026 Guide: Anthropic, Recursive, Sakana AI Show First Real Steps — TheTuringPost · 2026-09-18
- SceneAgent: agentic pipeline turns 3D captures into physics-ready scenes for robot training — hankyang94 · 2026-09-18
- Fari Research paper: misaligned AI may just persuade its human overseers — DG_Rand · 2026-09-18
- Biomedical world models: a framework for virtual drug trials, intervention design and planning — marinkazitnik · 2026-09-18
- LeanReact 0.1: expressing composable, provably correct React components in Lean — hargup13 · 2026-09-18
- Cell paper lays out how world models could transform biomedicine, from molecules to patients — marinkazitnik · 2026-09-18