One Epoch of Toloka's Enterprise RL Data Boosts Qwen3.5-27B Agent Benchmarks by up to 44pp

MParakhin · x · 2026-10-09

Toloka ran one epoch of RL post-training on Qwen3.5-27B using its off-the-shelf enterprise tool-use dataset, then evaluated on unseen agent benchmarks:

The takeaway: the model internalized enterprise workflow habits like "check the policy before changing anything," with the largest gains where training environments match the benchmarks. Full methodology and data are in Toloka's blog post.

Original post →

More from coding & agent

coding & agent channel →