BCE averaging convention was secretly tuning your learning rate
BlackHC · x · 2026-09-16
A subtle gotcha from the author's continual learning experiments: mean BCE averages over a pair's two outputs; switching to summed loss (×2) drops accuracy from 58.1% to 51.2% because at a fixed SGD learning rate it doubles the update. The averaging convention was doing implicit LR tuning — treat loss scaling as part of your optimizer setup.
Related event: BlackHC's toy continual learning experiment: PCA replay hits 91.1%(6 posts)→
More from Research
- Teaching a robotic hand to walk on its fingertips, no robot arm required — Scobleizer · 2026-09-16
- After searching thousands of learning rules, none beat backprop—and Nature confirms it's biologically plausible — aran_nayebi · 2026-09-16
- BoltzMol-1 finds WRN inhibitor hits for ~$10K in 10 days; best IC50 8.4 µM — GabriCorso · 2026-09-16
- Researchers clash over Sakana AI's bioplausible learning claim: MNIST results don't count — aran_nayebi · 2026-09-16
- A 0.62-AUC model made a 65,578-candidate materials search tractable, doubling hit rate — bravo_abad · 2026-09-16
- NeurIPS reference checker flags LLM-corrupted bib entries; desk rejection feared — suryanreddy · 2026-09-16