Salesforce turns its own agent config files into RL environments to train enterprise model Koa
omarsar0 · x · 2026-09-16
Salesforce trained its enterprise agent model Koa from open-weight Nemotron-3-Super-120B, using the declarative Agent Script files that define Agentforce agents as RL environments—expanding them into multi-turn tasks with simulated user personas, with rewards checking correct tool calls, trained via GRPO. Gains are modest but consistent: 69.41 on Tau2Bench (base: 68.64, GPT-4.1: 54.48), 0.86 on CRM Bench vs. 0.87 for Claude Opus 4.8, and function-call accuracy up from 0.71 to 0.77. The takeaway: structured workflow descriptions at your company may be convertible into RL environments.
More from coding & agent
- OpenAI reportedly orchestrates 10k internal agents without quality loss, while users' sub-agents 'just burn tokens' — RexDouglass · 2026-09-16
- Dev rebuilds Radish: a Redis server running entirely inside a Cloudflare Durable Object — whoiskatrin · 2026-09-16
- Agent looked right but wasn't: when do you stop double-checking AI outputs? — Luvena21 · 2026-09-16
- RSIAgent improves agents without weight updates, beats GPT-6 Astra on OSWorld 2.0 — Roger_M_Taylor · 2026-09-16
- Engineer uses LLMs to turn knowledge into playable simulations, builds a chip fab game — thehiphopswami · 2026-09-16
- Open-Source Skill Strips AI Slop: 41 Patterns Caught Across 4 Review Gates — Roger_M_Taylor · 2026-09-16