FLEET gives Best-of-N sampling memory via MCTS, hits GSM8K baseline with half the iterations
Helpful_Minimum_2214 · reddit · 2026-10-02
The authors of FLEET propose attributing external rewards to specific tokens, storing high-entropy states in a vector store with reward histories, and using a modified MCTS to adjust logits on the next run. On Llama 3.2 3B it matched the sampling baseline on GSM8K with half the iterations and lifted LiveCodeBench v6 easy from 0.59 to 0.69 under the same budget (9 vs 32 iterations). The method runs without sequential execution and the metadata store can serve as a prior for other tasks or SFT/RL. Paper and code are open.
More from Research
- Why AI can't solve the mystery of time: training presupposes the very clock it must explain — johnseach · 2026-10-03
- arXiv trends suggest AI uplift hits quantum physics within a year, all physics by late 2029 — cephaloform · 2026-10-03
- Berkeley's Humanoid Intelligence Center wins both tracks at IROS 2026 RoCo Challenge — berkeley_ai · 2026-10-03
- OpenStamp embeds watermarks into open-source LLM weights so users can't strip them — danish037 · 2026-10-03
- CruxBench lands at NeurIPS: frontier LLMs barely beat random at asking the right questions — mengyer · 2026-10-03
- Lampinen calls out neurosymbolic camp for walking back internal-symbolism claims — AndrewLampinen · 2026-10-03