Backprop Through Latents Beats Discrete Thinking for Credit Assignment

ShikharMurty · x · 2026-08-26

In a discussion contrasting discrete vs. latent thinking, the author argues that backpropagating through latent thoughts is more informative for credit assignment: you get the full gradient of log π(good action) with respect to the latents — an exact signal for how each latent should change. With discrete thought tokens, there is no derivative telling the model how thought token 50 should have been different. This refines his earlier claim that RL over sampled discrete actions works the same either way.

Original post →

More from Research

Research channel →