Repeated sampling lifts DeepSeek-Coder on SWE-bench Lite from 15.9% to 56%
le_james94 · x · 2026-09-16
- DeepSeek-Coder-V2-Instruct solves 15.9% of SWE-bench Lite issues with 1 sample, but 56% with 250 samples.
- The trick works because a test suite tells you which candidate was right — when the verifier is free, repeated sampling hugely amplifies generation.
Related event: Three Weeks Through Stanford CS329A: Generators Have Outrun Verifiers(9 posts)→
More from Research
- Humans scale superlinearly on 14-day tasks while agents hit log-linear limits — AlexGDimakis · 2026-09-16
- Old World journal launches third best-article competition for AI-driven humanities research — ArtificialOther · 2026-09-16
- Hugging Face releases a beginner-friendly visual guide to flow matching — Hugging Face · 2026-09-16
- GPT Astra improves Anthropic's zeta zeroes certificate from 67.25% to 67.31% — dsanft · 2026-09-16
- Scott Aaronson: the AI skeptics' twenty-year-old defensible position has collapsed — avt_im · 2026-09-16
- Google's 53.9M-param diffusion model speeds search query expansion 12-20x — imjustnewatai · 2026-09-16