Prime Values: Open-Source Value Function Infrastructure
_AndrewZhao · x · 2026-07-19
This is an introduction to an open-source training infrastructure around value functions. The author says that frontier labs are widely using value functions, but the open ecosystem lacks good infrastructure, so they built Prime Values on top of prime-rl.
Key features include:
- Clean, separable, and hackable abstraction layers
- Asynchronous value trainer/evaluator nodes to avoid trainer bottlenecks
- Native value warmup support
- Streaming replay buffer leveraging value model's higher tolerance for 'older data/reuse'
The author also claims that the default configuration can match or outperform mean-baseline GRPO on single-turn and multi-turn tasks.
More from coding & agent
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11