Tropical Reinforcement Learning: Summing Probabilities Can't Teach LLMs Composition

burny_tech · x · 2026-10-07

Researchers present Tropical Reinforcement Learning, arguing standard RL fails to teach LLMs composition for a fundamental reason: its algebra is suboptimal. Summing the probabilities of all successful answers measures how often a model succeeds, not which solution succeeds; a max-plus (tropical) algebra instead tracks the best solution path. A foundational critique and new framework for RL on LLMs.

Original post →

More from Research

Research channel →