Researcher uses Claude Code to build a tiny 10k tok/s CPU neural net, loss still descending

GregoryDiamos · x · 2026-09-07

Gregory Diamos argues we should revisit outrageously small neural networks. Needing a CPU model running at 10k tok/s for data processing, he gave Anthropic's Claude Code a pile of tokens to build one. Across a 4.91B-token run, smoothed training loss fell monotonically within each curriculum phase and was still descending at the end, with three interesting discoveries along the way.

Related event: Developer Uses Claude Code to Build Tiny CPU Model Hitting 10k+ tok/s(4 posts)→

Original post →

More from coding & agent

coding & agent channel →