From-Scratch Model Trained with $1000 Reaches 11% on SWE-bench
sanmikoyejo · x · 2026-08-14
A new study from the Max Planck Institute and Stanford explores the limits of 'speedrunning' SWE-bench on a shoestring budget.
- Setup: Built upon Andrej Karpathy's nanochat, the model is trained from scratch on Qwen Coder agentic trajectories, with SWE-bench codebases strictly held out.
- Surprising Results: A mere $10 in compute yields a non-zero solve rate (0.53%). For $60 (approx. 12 B200-hours), the model achieves 5.0% pass@1, matching the SOTA at the benchmark's release (Claude 2).
- Scaling: Solve rate improves log-linearly with compute, reaching 11.0% at the $1000 mark—landing between Claude 3 Haiku and Claude 3 Opus.
The authors caution that climbing the benchmark without gains in underlying knowledge or math capabilities measures something narrower than general capability.
Related event: Researchers Train SWE-bench Model for $60, Rivaling Claude 2(2 posts)→
More from coding & agent
- Google Joins as Core Maintainer of Agent Plugins 1.0.0 to Standardize Skills — jggomezt · 2026-08-14
- Nuphos Launches AI-Native DevOps Workspace for Safe Agent Interventions — Aiden_Tech_Ai · 2026-08-14
- Meta Launches 'Muse Code' Terminal Agent with Replayable Runtime Architecture — JeremyCMorgan · 2026-08-14
- proxmux: Open-Source Terminal Dashboard for Managing Proxmox VE — tom_doerr · 2026-08-14
- AI Marketing Tool Combining MCP to Auto-Scrape and Categorize Videos — eptwts · 2026-08-14
- AI Coding Destroys Developer Apprenticeship: Where Will Novices Learn? — omojumiller · 2026-08-14