Kimi-style agentic reward modeling adds rubric scoring and budgeted verbosity control
stochasticchasm · x · 2026-07-28
The discussion focuses on an Agentic Generative Reward Model (GRM) for non-verifiable general tasks.
- It keeps tournament-style group rewards with binary comparisons, as in Kimi K2.5.
- For better judgment, the agentic judge must follow a fixed protocol: read the output, generate a rubric, score each candidate against that rubric, and record the scores.
- To reduce reward hacking from overly verbose answers, the method adds budget-based verbosity control: an initial verbosity budget is estimated from a cold-start model, then scaled by a multiplier, and outputs that exceed the budget automatically lose the comparison.
The reply adds that the approach is especially interesting for reasoning-effort allocation, notes extensive SFT, and praises the choice to estimate the initial token budget from the post-SFT model and scale it with a multiplier.
Related event: New Paper Proposes Agentic Reward Model and RL Inference Budget Control(2 posts)→
More from coding & agent
- Amp orbs go from 2.6% to 97.5% of weekly credits in five weeks — glenbeer · 2026-07-28
- Agents are already accelerating research in a knowledge-graph task synthesis system — stochasticchasm · 2026-07-28
- xAI schedules a 12-hour Grokathon in San Francisco for August 8 — shaunmmaguire · 2026-07-28
- A unified RL harness can swap in Kimi Code, Claude Code, Codex, and more — stochasticchasm · 2026-07-28
- ChatGPT Work users are automating ticket alerts, listings, and trip planning from their phones — OpenAIDevs · 2026-07-28
- Open-source repo bundles 129 practical AI app, agent, and RAG projects — Arindam_1729 · 2026-07-28