Training a Model for $60 Beats Claude 2 on SWE-bench

StanfordAILab · x · 2026-08-14

A developer turned Karpathy's nanochat into a SWE-bench speedrun, demonstrating that high-performance coding models can be trained at extremely low costs.

Experiment results:

Starting from randomly initialized model weights, this experiment proves that with an optimized training pipeline, the coding capabilities of early frontier models can be replicated or even surpassed at a fraction of the compute cost.

Original post →

More from Research

Research channel →