Prime Intellect brings multi-agent training to its open RL stack PRIME-RL
willcb · x · 2026-09-29
Prime Intellect announced that its open-source RL stack PRIME-RL now supports multi-agent system training, moving beyond single-agent RL. The update lets you program arbitrary agent interactions, choose which roles learn, and assign credit across the full interaction.
- Built on two abstractions: Agent (wrapping Taskset/Harness/Runtime, with signature Agent.run(task) -> Trace) and Env
- New expressible patterns: Agentic Judging (a judge grades solver traces), Self-Play, and User-Sim (a user agent interacting with an assistant)
- Vincent Weisser frames the broader roadmap as "superintelligence emerging from trillions of self-improving agents rather than one god model," spanning autonomous AI research, a recursively self-improving agent harness, and continual learning
A substantive expansion of agentic RL infrastructure, fully open source.
More from coding & agent
- Claude Sonnet 5.5 lands in Claude Code: 30% faster, up to 20% lower platform costs — lydiahallie · 2026-09-29
- marimo adds Pixi sandbox support for portable notebooks with CUDA, R, ffmpeg — S_Conradi · 2026-09-29
- Anthropic's guidance: start with Opus for large projects; max-effort Sonnet loses its balance — ClaudeDevs · 2026-09-29
- VSArena V1 ships: evaluate embodied/VLA agents without physical robots — NovaCoding · 2026-09-29
- NVIDIA introduces Open Agent Security Platform to boost AI agent safety — HumanSoulAI · 2026-09-29
- HeyGen puts Jev in front of its MCP to filter leads in milliseconds for fractions of a cent — HeyGen · 2026-09-29