Marathoner: a 9B model that codes for 10+ hours, lifting SWE-bench Verified to 77.5%
mark_k · x · 2026-10-01
Marathoner: persistent coding in a 9B model
- Researchers from Ant Group, Peking University and University of Macau report Marathoner, built on Qwen3.5-9B, capable of 10+ hours of continuous coding and 1,000+ tool calls.
- Training: software tasks mined from 100K GitHub PRs across 10K repos, chained for difficulty; trained on successful Kimi K3 runs; RL in real execution environments; a bonus specifically rewards useful late-run progress like finding hidden bugs.
- Results: SWE-bench Verified 43.8% → 77.5%; Terminal-Bench 2.0 27.3% → 57.2%. Still behind frontier models, but a huge gain from the same 9B start.
- Author's take: an agent that keeps making progress while you sleep changes what you can delegate — and this persistence can be trained into small models.
Related event: Ant Group Unveils Marathoner, an Agent That Codes for 10+ Hours Straight(2 posts)→
More from coding & agent
- One prompt, 3 hours, 22.8M tokens: local quantized model builds a GTA-style game — zmarcoz2 · 2026-10-01
- Building a million-page OCR pipeline with a 500GB RAM used server plus LLM extraction — oilmutt · 2026-10-01
- Agentic coding lowered the cost of change and raised the cost of holding invariants — _AustinCalvert_ · 2026-10-01
- Developer wakes up to 6.5 hours of agent-completed work — calls it 2 weeks of his time — therealdanvega · 2026-10-01
- Jev vs PydanticAI: control the decision before the model responds, with confidence scores — KhuyenTran16 · 2026-10-01
- Ex-OpenAI Researcher Launches Jev: A System One Model for Structured Decisions, 100x Faster — KhuyenTran16 · 2026-10-01