IBM releases DRACO: dynamic rubrics give per-step credit assignment for verifier-free agent RL
ibm-research · hf · 2026-09-04
- IBM Research released DRACO on Hugging Face for long-horizon agent reinforcement learning.
- Method: it dynamically generates rubrics and redistributes trajectory-level scores into per-step advantages, enabling credit assignment without verifiers.
- Claimed to improve long-horizon agent performance in RL training.
More from coding & agent
- LukeW's Intent update: multi-agent coordination, workspaces, multi-device runs — LukeW · 2026-09-04
- Intent update coordinates swarms of agents with workspaces across devices — LukeW · 2026-09-04
- VC advice for agentic startups: pitch how your agent recovers from failures — atShruti · 2026-09-04
- WebMCP demo: an AI agent DJs in the browser with hundreds of direct tool calls — No_Guide_8697 · 2026-09-04
- Newsletters, not killer use cases, are what get users to trust agents with real email access — JanJanJaJa · 2026-09-04
- Kafka, Kafka Connect and Schema Registry exposed as native MCP tools — jkriket · 2026-09-04