Three weeks, 19 lectures: a deep recap of Stanford AA203 from Euler equation to PPO

le_james94 · x · 2026-10-09

Engineer James Le spent three weeks working through all 19 lectures of Stanford's Spring 2026 AA203 (Optimal and Learning-Based Control, taught by Marco Pavone and Daniele Gammelli) and published a lecture-by-lecture recap. The course rebuilds one toy particle-control problem repeatedly: hand-derived Hamiltonian and costates, then numerical solvers, finally neural policies trained on rollouts once the model is gone. He rejects the chronology reading that PPO simply replaces Pontryagin—Lecture 19's repairs are Lectures 11 and 12 in new clothes—and offers a sharp critique of model-based planning: a planner hunting high predicted reward finds exactly where the model errs in a flattering direction.

Related event: "Model proposes, system executes": guardrails for agent tool calls(3 posts)→

Original post →

More from Research

Research channel →