Planning against a learned model seeks out exactly where the model errs flatteringly
le_james94 · x · 2026-10-09
In a RL derivation thread, lejames94 argues learned-model planning only works for linear systems: a planner chasing high predicted reward finds exactly where the model errs in a flattering direction — the optimizer seeks the error out. He also derives REINFORCE: expanding log pθ(τ), initial-state and dynamics terms vanish, so the policy gradient is the max-likelihood gradient weighted by reward.
Related event: "Model proposes, system executes": guardrails for agent tool calls(3 posts)→
More from Research
- Exa launches ATLAS benchmark: even priciest search agents miss ~1/3 of results — yoimnotkesku · 2026-10-09
- PPTBench: a new benchmark testing if coding agents can rebuild visuals into editable slides — jiqizhixin · 2026-10-09
- TIDE attributes diffusion outputs to training images in milliseconds — serrjoa · 2026-10-09
- DeepScholar-Bench at COLM 2026: benchmarking AI-generated research synthesis — mrdrozdov · 2026-10-09
- Frontier AI models beat human experts at earnings predictions for the first time — maithra_raghu · 2026-10-09
- Cell paper reconstructs cell fate map of the mouse embryo — anshulkundaje · 2026-10-09