Tropical Reinforcement Learning: Summing Probabilities Can't Teach LLMs Composition
burny_tech · x · 2026-10-07
Researchers present Tropical Reinforcement Learning, arguing standard RL fails to teach LLMs composition for a fundamental reason: its algebra is suboptimal. Summing the probabilities of all successful answers measures how often a model succeeds, not which solution succeeds; a max-plus (tropical) algebra instead tracks the best solution path. A foundational critique and new framework for RL on LLMs.
More from Research
- eigenrobot: automating math papers is easy, and most of economics and theoretical physics is next — eigenrobot · 2026-10-07
- François Fleuret: math is unique in that its truths are "context free" — francoisfleuret · 2026-10-07
- Two Years After First Reasoning Model, AI Has Produced '20 Fields Medals' of New Math — __nmca__ · 2026-10-07
- NUS releases SafeActBench: 656 cases reveal where tool-using agents break the evidence-to-action chain — NationalUniversityofSingapore · 2026-10-07
- MEND: RL for flow models via proximal velocity matching beats Flow-GRPO in 100 vs ~4k updates — UTEXAS · 2026-10-07
- JLD: perceptual distance from a frozen encoder's Jacobian, fitted in 35s from 100 images, beats LPIPS and DISTS — Shreshth Saini · 2026-10-07