VPO: A Training Method to Boost Solution Diversity
bronzeagepapi · x · 2026-07-19
Vector Policy Optimization (VPO) is a training method designed to mitigate the issue of "declining solution diversity," a problem that limits the gains achieved during test-time search.
The training pipeline roughly works as follows: the model sequentially generates multiple candidate answers, which are evaluated using multiple objectives and randomized reward weights. The highest-scoring solution for each set of weights is then selected and averaged to produce a scalar reward. The author notes that this reward calculation method can be combined with GRPO.
More from Research
- Statistical theory paper studies how fast signatures learn in path regression — chaumian · 2026-07-21
- PROWL uses a world model to keep Minecraft agents exploring after failures — nathanbenaich · 2026-07-21
- LeRobot v0.6.0 adds end-to-end 3D depth training data for robots — RemiCadene · 2026-07-21
- Multiagent v2 playbook calls for 64 agents, diverse proof routes and adversarial checks — danshipper · 2026-07-21
- HarmonicMath says Lean autonomously solved eight previously studied open problems — MarioKrenn6240 · 2026-07-21
- SeeSE3 finds 3D structure emerging in frozen vision features and camera-pose alignment — ducha_aiki · 2026-07-21