Functional Gradient Descent Beats Neural Nets by ~10x, New Paper Fixes Its Convergence Flaw

burkov · x · 2026-10-01

Andriy Burkov highlights a new paper arguing functional gradient descent algorithms generally outperform neural nets—often by an order of magnitude—but are hard to implement because functional gradients are infinite-dimensional and naive approximations converge to the wrong place.

The paper formalizes a broad class of approximation schemes ("adaptive representations") that provably guarantee convergence to the global minimizer while being immediately implementable, with the resulting algorithms beating corresponding neural nets across many settings.

Original post →

More from Research

Research channel →