Training a Model for $60 Beats Claude 2 on SWE-bench
StanfordAILab · x · 2026-08-14
A developer turned Karpathy's nanochat into a SWE-bench speedrun, demonstrating that high-performance coding models can be trained at extremely low costs.
Experiment results:
- $60 in compute: Achieved a 5.0% pass@1 rate on SWE-bench, comparable to the 2023 SOTA model Claude 2.
- $1000 in compute: Achieved an 11.0% pass@1 rate, surpassing Claude 3 Haiku.
Starting from randomly initialized model weights, this experiment proves that with an optimized training pipeline, the coding capabilities of early frontier models can be replicated or even surpassed at a fraction of the compute cost.
More from Research
- Reproducing 2,200 ICML Papers with Agents Reveals Falsifications — QGallouedec · 2026-08-14
- Developer Compiles Doom's Rendering Algorithm Directly into Transformer Weights — notforrob · 2026-08-14
- Prototype Shows Qwen4B Updating Memory in Real-Time Without Retraining — SpearHammer · 2026-08-14
- LLMs Are Not Stateless: Paper Reveals Implicit Memory Threatens Agent Eval Safety — lbeurerkellner · 2026-08-14
- Frontier LLMs Hit Perfect Detection Rate in UEFI Firmware Vulnerability Tests — evilsocket · 2026-08-14
- Personalized T-Cell Therapy Eliminates Metastatic Cancer in Teenager — Dr_Singularity · 2026-08-14