Salesforce turns its own agent config files into RL environments to train enterprise model Koa

omarsar0 · x · 2026-09-16

Salesforce trained its enterprise agent model Koa from open-weight Nemotron-3-Super-120B, using the declarative Agent Script files that define Agentforce agents as RL environments—expanding them into multi-turn tasks with simulated user personas, with rewards checking correct tool calls, trained via GRPO. Gains are modest but consistent: 69.41 on Tau2Bench (base: 68.64, GPT-4.1: 54.48), 0.86 on CRM Bench vs. 0.87 for Claude Opus 4.8, and function-call accuracy up from 0.71 to 0.77. The takeaway: structured workflow descriptions at your company may be convertible into RL environments.

Original post →

More from coding & agent

coding & agent channel →