Just 500 curated tasks and 15 GRPO steps lifted Qwen3.8-27B by 11.3 points

ycombinator · x · 2026-10-09

AfterQuery, an AI training-data company, shared a training experiment: using only 500 tasks from its SWE agent dataset, its internal post-training team improved Qwen3.8-27B-Medium by 11.3 points in just 15 steps of GRPO, arguing that not all coding data is created equal.

Its co-founder added that building an in-house research team was one of the company's best early decisions: knowing what great data looks like is table stakes, and the real value is showing labs where their model's losses are and supplying data to close those gaps.

Related event: Fine-Tuning on Just 500 Tasks Boosts Qwen 27B by 11.3 Points(2 posts)→

Original post →

More from coding & agent

coding & agent channel →