Salesforce Paper: Only 35.8% of the Cleanest Public RL Environments Pass Audit

dair_ai · x · 2026-09-24

A Salesforce AI Research paper highlights the importance of good verifiers in RL environments: an audit found only 35.8% of environments in the cleanest public RL collection for terminal agents were sound, with two other collections at just 10.1% and 3.3%.

Key findings:

Original post →

More from coding & agent

coding & agent channel →