500 curated SWE tasks lift Qwen 27B by 11.3 points in 15 GRPO steps

ycombinator · x · 2026-10-09

AfterQuery reports that post-training Qwen3.8-27B-Medium on just 500 tasks from its SWE agent dataset improved the model by 11.3 points within 15 GRPO steps. The takeaway: curated coding data beats volume — small, high-quality task sets can deliver outsized gains. Full training experiment details are in their blog post.

Related event: Fine-Tuning on Just 500 Tasks Boosts Qwen 27B by 11.3 Points(2 posts)→

Original post →

More from Research

Research channel →