DiVeR preprint trains VLA verifiers only on decision-critical states for test-time scaling
SharonYixuanLi · x · 2026-10-07
- New preprint DiVeR (Decision-Critical Verifier Learning) studies how to learn better verifiers for Vision-Language-Action (VLA) models, improving the Generate → Verify → Act loop via Best-of-N action selection.
- Core observation: not all states matter equally. During routine motion (traversing free space, carrying objects) candidate actions are similar, offering little signal; at decision-critical moments like grasp or placement alignment, candidates diverge and small differences change outcomes.
- The method identifies decision-critical states by measuring disagreement among sampled actions and trains verifiers on them. The work is tied to the NeurIPS 2026 PTA workshop (From Pretrained Representations to Acting Agents), with invited speakers including Aviral Kumar, Benjamin Eysenbach, and Sherry Yang.
More from Embodied
- Indian builders make offline AI hardware that listens and remembers without apps — alysha_lobo · 2026-10-07
- Generalist's GEN-1.5 robot model installs and removes O-rings autonomously, learns from seconds of demos — 141_1337 · 2026-10-07
- YC-backed Vexo launches always-listening AI bracelet that keeps your context — ycombinator · 2026-10-07
- H-JEPA: Hierarchical world models lift visual AntMaze success from 18% to 73% — ylecun · 2026-10-07
- Million-mile trucker after Tesla Semi factory tour: this industry is going to be changed — elonmusk · 2026-10-07
- Attacca trains embodied agents for state-continuity, up to 7x on long-horizon tasks — nanyang-technological-university-singapore · 2026-10-07