Just 500 curated tasks and 15 GRPO steps lifted Qwen3.8-27B by 11.3 points
ycombinator · x · 2026-10-09
AfterQuery, an AI training-data company, shared a training experiment: using only 500 tasks from its SWE agent dataset, its internal post-training team improved Qwen3.8-27B-Medium by 11.3 points in just 15 steps of GRPO, arguing that not all coding data is created equal.
Its co-founder added that building an in-house research team was one of the company's best early decisions: knowing what great data looks like is table stakes, and the real value is showing labs where their model's losses are and supplying data to close those gaps.
Related event: Fine-Tuning on Just 500 Tasks Boosts Qwen 27B by 11.3 Points(2 posts)→
More from coding & agent
- LLM-as-a-Verifier: Weaker Model Verifies Stronger One, Hits 69.2% SOTA on Terminal-Bench 4 — Azaliamirh · 2026-10-09
- Salesforce's SRD distills hindsight into foresight, lifting 2B agent success from 0% to 60.6% — Salesforce · 2026-10-09
- Agent access problem measured: 7 of 16 subreddits blocked an agent before it posted — lulzxdxdxd · 2026-10-09
- Prompt Tuning Is Forgotten Lore — Are We Massively Underusing Finetuned Tokens? — cephaloform · 2026-10-09
- System architect shares WhatsApp voice-note AI assistant setup — No_Kangaroo_4454 · 2026-10-09
- AI Workspace Rolls Out Multiplayer, Interactive Dashboards, Memory & Skills — heyneighbor · 2026-10-09