Pairwise BCE lifts continual-learning accuracy from 19.4% to 58.1%, and Adam forgets more than SGD

BlackHC · x · 2026-09-16

The author ran an automated research experiment (done with Claude) on a toy continual-learning setup: a small MLP learns Split MNIST sequentially (0/1 → 2/3 → … → 8/9), each stage using only that pair's training images, then gets tested on all ten digits (averaged over 3 seeds).

Related event: BlackHC's toy continual learning experiment: PCA replay hits 91.1%(6 posts)→

Original post →

More from Research

Research channel →