nanoswe hits 15.7% on SWE-bench with $1000 of compute, matching Claude 3 Opus
sanmikoyejo · x · 2026-10-03
The nanoswe community speedrun set a new record: 11.8% → 15.7% on SWE-bench Verified, matching Claude 3 Opus (15.8%), trained from scratch with roughly $1,000 of compute (192 B200-hours).
- $60 track (12 B200-hours): nanochat + long context + qwen coder trajectories, 5.0%
- $1000 track: adding nemotron lightning trajectories pushed it to 15.7%; earlier milestones were 11.0% (base) and 11.8% (fixed vLLM serving)
- Targets to beat: GPT-4o 23.2%, Claude 3.5 Sonnet 33.6%
- Weights, evals, and training scripts are open; anyone can submit improved runs
More from coding & agent
- Every business will have its own agent the way it has a website, argues Every — every · 2026-10-03
- AI Conference Day 2: Agent Inference Costs and Observability Steal the Show — sanjaykalra · 2026-10-03
- Open-Source Repo Ships 28 Portable Agent Skills for Structured Reasoning Across Coding Agents — tom_doerr · 2026-10-03
- Sign in with ChatGPT lets Plus/Pro users spend plan usage across 11 partner apps — _AustinCalvert_ · 2026-10-03
- Best tech stacks for agentic coding: training data volume decides everything — aiblastoff · 2026-10-03
- After auditing dozens of companies: the 4 reasons custom AI workflows fail in production — salespire · 2026-10-03