FLEET gives Best-of-N sampling memory via MCTS, hits GSM8K baseline with half the iterations

Helpful_Minimum_2214 · reddit · 2026-10-02

The authors of FLEET propose attributing external rewards to specific tokens, storing high-entropy states in a vector store with reward histories, and using a modified MCTS to adjust logits on the next run. On Llama 3.2 3B it matched the sampling baseline on GSM8K with half the iterations and lifted LiveCodeBench v6 easy from 0.59 to 0.69 under the same budget (9 vs 32 iterations). The method runs without sequential execution and the metadata store can serve as a prior for other tasks or SFT/RL. Paper and code are open.

Original post →

More from Research

Research channel →