NVIDIA's Skill2Env turns 3.4k Agent Skills into 8k RL environments, boosting Qwen-27B by 4.7 points
burny_tech · x · 2026-09-23
NVIDIA researchers released Reinforcing Agents with Collective Skills on alphaXiv, introducing Skill2Env, a human-aligned data recipe for modern agentic RL.
- Idea: public Agent Skills encode domain expertise, reusable workflows and success metrics tied to real-world tasks. A pipeline converts 3,400 web-crawled, filtered Agent Skills into 8,000 executable terminal environments, each scored by both programmatic tests and behavioral rubrics.
- Results: after just 300 RL steps with a DPPO variant, Qwen-3.8 27B gains 4.7 percentage points on Terminal-Bench 2.1 and 4.3 points on S2EBench, a hand-verified private benchmark of real agent use cases.
- Analysis: models show increased behavioral alignment with the source Agent Skills on similar problems, framed as a path toward value-aligned superintelligence.
More from coding & agent
- One vanilla JS shader, no assets: AI agent renders a golden-hour sea-cliff lighthouse — NathanWilbanks_ · 2026-09-23
- Fixing agent memory: InfoWorld argues AI memory must be inspectable to be trustworthy — rseroter · 2026-09-23
- 27k-tool x402/MPP catalog released, no longer gatekept — MurrLincoln · 2026-09-23
- A week with Claude Opus 5.5: six tips to save your credits — every · 2026-09-23
- ML author Burkov slams OpenCode as 'one huge bug' — burkov · 2026-09-23
- Dev joining Runway shares the homemade AI bootcamp resources he used to get up to speed — tlakomy · 2026-09-23