500 curated SWE tasks lift Qwen 27B by 11.3 points in 15 GRPO steps
ycombinator · x · 2026-10-09
AfterQuery reports that post-training Qwen3.8-27B-Medium on just 500 tasks from its SWE agent dataset improved the model by 11.3 points within 15 GRPO steps. The takeaway: curated coding data beats volume — small, high-quality task sets can deliver outsized gains. Full training experiment details are in their blog post.
Related event: Fine-Tuning on Just 500 Tasks Boosts Qwen 27B by 11.3 Points(2 posts)→
More from Research
- DNA Typewriter reconstructs mouse embryo lineage: 1.34M profiled cells from zygote to E13.5 — anshulkundaje · 2026-10-09
- Yacine calls for a code reuse benchmark to hill-climb model behavior — yacinelearning · 2026-10-09
- Data companies are becoming research labs, with verifier design the most climbable problem — madhavsinghal_ · 2026-10-09
- Amazon runs Karpathy's AutoResearch at production scale for 12 weeks, finds 5 failure modes — amazon · 2026-10-09
- ReGain: training-free fix restores subject fidelity lost from personalizing on synthetic images — UIUC-CS · 2026-10-09
- OpenAI's claimed Navier–Stokes solution reportedly doesn't match its Lean verification — kyan100 · 2026-10-09