DeepSWE-mini: a 16-instance subset that replicates the full DeepSWE leaderboard rankings

asankhs · reddit · 2026-09-18

A new dataset, DeepSWE-mini (on Hugging Face: LocalLLaMA/deepswe-mini), is a 16-instance subset of the DeepSWE coding benchmark. Analysis showed it faithfully replicates the relative rankings (not absolute scores) of the full benchmark, letting developers quickly benchmark new local models without the full, slow run.

Related event: DeepSWE-mini Released: 16 Instances Reproduce Full Leaderboard(2 posts)→

Original post →

More from Research

Research channel →