Ant Group's Marathoner Trains Agents to Work 10+ Hours With 1000+ Tool Calls
antgroup · hf · 2026-09-30
Ant Group introduces Marathoner, an autonomous agentic model for ultra-long-horizon execution. Its post-training pipeline synthesizes tasks from major GitHub release PRs (1000+ lines), chains tasks for frontier difficulty, applies rejection-sampling finetuning with teacher-generated trajectories, and runs RL in real sandboxes with a Later Stage Bonus Reward. On 5 ultra-long-horizon benchmarks it consistently beats its base model and strong proprietary models, sustaining 10+ hours of work and 1000+ tool calls.
More from coding & agent
- Cube Launches Always-On Cloud Computers for Running Claude Code and Codex Agents — algo_diver · 2026-09-30
- A new auto-research loop that bootstraps the shape of the best possible result — burny_tech · 2026-09-30
- Stack Overflow joins OpenAI DevDay to share how Codex sped up its new architecture — pchandrasekar · 2026-09-30
- Models improving doesn't obsolete your agentic coding scaffolding, argues pushback on viral take — max_paperclips · 2026-09-30
- Building a Code Review Agent That Learns From Feedback With Groq and Hindsight — pasulabhavya · 2026-09-30
- Open-Dots, an open-source clone of OpenAI's Dots, hits 4,500 GitHub stars in 24 hours — matchaman11 · 2026-09-30