One Epoch of Toloka's Enterprise RL Data Boosts Qwen3.5-27B Agent Benchmarks by up to 44pp
MParakhin · x · 2026-10-09
Toloka ran one epoch of RL post-training on Qwen3.5-27B using its off-the-shelf enterprise tool-use dataset, then evaluated on unseen agent benchmarks:
- External benchmarks: τ³ retail 77.6% → 86.8% (+8.9pp, 95% CI [+4.5, +13.6]); AutomationBench Ops partial credit +9.6pp, Support +7.6pp; Toolathlon +4.3pp
- Cross-domain transfer: on an entirely held-out OTS environment, pass rate jumped 14.0% → 58.0% (+44.0pp)
- Held-out OTS tasks: pass rate 41.6% → 64.5% (+22.8pp); tasks passed on all three attempts doubled from 22.9% to 46.4%
- Reward design: at equal steps, a rubric-based reward beat binary pass/fail by +9.4pp
The takeaway: the model internalized enterprise workflow habits like "check the policy before changing anything," with the largest gains where training environments match the benchmarks. Full methodology and data are in Toloka's blog post.
More from coding & agent
- Temporal runs the Pi coding agent on durable execution to survive machine failures — francesc · 2026-10-09
- Zed CEO Nathan Soller: AI-generated unit tests are slop, integration tests are the right middle ground — zeeg · 2026-10-09
- Jev pitched as the fastest AI model for agents: millisecond decisions at near-zero cost — Arindam_1729 · 2026-10-09
- Devs debate whether LLMs should write tests: 'tests expose things memory can't hold' — ivan_bezdomny · 2026-10-09
- Musk shows Grok agent setting up 16 emulators on a handheld with one prompt — elonmusk · 2026-10-09
- MCPtoAI: open-source client shares MCP tools across AI models while credentials stay on-device — abktya · 2026-10-09