Researcher uses Claude Code to build a tiny 10k tok/s CPU neural net, loss still descending
GregoryDiamos · x · 2026-09-07
Gregory Diamos argues we should revisit outrageously small neural networks. Needing a CPU model running at 10k tok/s for data processing, he gave Anthropic's Claude Code a pile of tokens to build one. Across a 4.91B-token run, smoothed training loss fell monotonically within each curriculum phase and was still descending at the end, with three interesting discoveries along the way.
Related event: Developer Uses Claude Code to Build Tiny CPU Model Hitting 10k+ tok/s(4 posts)→
More from coding & agent
- Agent memory is not the context layer, and 'enterprise agent memory' may be a fake category — femke_plantinga · 2026-09-07
- Dev ships two iOS apps and an Obsidian plugin in one go with Codex — vista8 · 2026-09-07
- Korea's top game studio opens its animation/sound assets: VARCO 3D MCP and VARCO Sound — arrakis_ai · 2026-09-07
- Qwen Code ships cua-driver v0.20.4 with signed macOS, Linux and Windows binaries — github-actions[bot] · 2026-09-07
- Spotify cut Claude Code token usage by 90% by routing big file reads to cheap models — aliscodes · 2026-09-07
- Knowledge work is much harder for agents to automate than code, argues Matt Pocock — mattpocockuk · 2026-09-07