Functional Gradient Descent Beats Neural Nets by ~10x, New Paper Fixes Its Convergence Flaw
burkov · x · 2026-10-01
Andriy Burkov highlights a new paper arguing functional gradient descent algorithms generally outperform neural nets—often by an order of magnitude—but are hard to implement because functional gradients are infinite-dimensional and naive approximations converge to the wrong place.
The paper formalizes a broad class of approximation schemes ("adaptive representations") that provably guarantee convergence to the global minimizer while being immediately implementable, with the resulting algorithms beating corresponding neural nets across many settings.
More from Research
- New lower bounds for kissing numbers: τ₁₉≥12268 and τ₂₁≥30761, open-sourced — felpix_ · 2026-10-01
- New paper: LLMs can get answers right while their chain-of-thought traces are invalid — rao2z · 2026-10-01
- Gemini 3.8 Flash (high) hits 84.8% on WeirdML v2, first Flash to beat Gemini 3.1 Pro — teortaxesTex · 2026-10-01
- Researchers steal frontier models' hidden reasoning by replaying encrypted chain-of-thought traces — maksym_andr · 2026-10-01
- Thinking in Geometric Terms: What ReLU, LayerNorm, LoRA and VQ Do to the Data Space — techNmak · 2026-10-01
- First superhuman Stratego AI unveiled in Nature paper using RL and test-time compute — zicokolter · 2026-10-01