Blog Series Revisits AdaGrad, Reproducing Full Derivation via Upper Bound Minimization
aaron_defazio · x · 2026-09-05
Part IX of the 'Revisiting Convergence Results in Convex Optimization' series takes a fresh look at AdaGrad, the seminal adaptive gradient algorithm. The post reproduces its full derivation via a 'convergence analysis → minimize the upper bound → optimal preconditioning matrix' pipeline, placing adaptive learning-rate methods within a unified convex-optimization convergence framework.
More from Research
- Reverse-engineering how networks encode symbolic structure as compositional vectors — tallinzen · 2026-09-06
- Meta's autonomous research agent AIRA3 wins Kaggle gold, placing 8th of ~4,000 teams — andrew_n_carr · 2026-09-06
- Xbench turns Twitter into an AI eval, tracking real sentiment and model switches — thedealdirector · 2026-09-06
- Principia: A Benchmark Testing Whether Video Models Grasp Physics via Pendulum Dependencies — anand_bhattad · 2026-09-06
- Jerry Liu: The Other Half of RLMs Is Programmatic Recursion, aka 'Dynamic Workflows' — lateinteraction · 2026-09-06
- Huawei & CUHK open-source Lego-RL: plug real coding agent harnesses into RL, SWE-bench 64.0→70.4 — 青稞AI · 2026-09-06