LenVM lands COLM 2026 Spotlight: value model predicts remaining generation length for efficient reasoning

xwang_lk · x · 2026-10-06

LenVM (Length Value Model) will be presented as a Spotlight paper at the Efficient Reasoning Workshop of COLM 2026 (Oct 5–9), with a main-conference poster on Oct 7 (#42).

Core idea: tokens are the basic unit of inference compute; length drives cost, latency, KV cache and reasoning quality, and token budgets become a bottleneck as reasoning chains and agentic workflows grow. Yet length is rarely modeled systematically — most methods operate only at the coarse sequence level.

LenVM connects length modeling with reward/value modeling: assign a constant cost to every generated token, and remaining length becomes a value prediction problem — yielding dense, unbiased, annotation-free supervision and a new dimension of scaling for length modeling.

Original post →

More from Models

Models channel →