Planning against a learned model seeks out exactly where the model errs flatteringly

le_james94 · x · 2026-10-09

In a RL derivation thread, lejames94 argues learned-model planning only works for linear systems: a planner chasing high predicted reward finds exactly where the model errs in a flattering direction — the optimizer seeks the error out. He also derives REINFORCE: expanding log pθ(τ), initial-state and dynamics terms vanish, so the policy gradient is the max-likelihood gradient weighted by reward.

Related event: "Model proposes, system executes": guardrails for agent tool calls(3 posts)→

Original post →

More from Research

Research channel →