SWE Odyssey benchmark tests long-horizon autonomous agents

garrytan · x · 2026-08-15

SWE Odyssey is a new standardized benchmark designed to address the saturation of existing benchmarks. It aims to quantifiably test a larger claim: can agents work autonomously for hours and still build the correct thing? It is part of a family of ultra-long-horizon evaluations.

Original post →

More from coding & agent

coding & agent channel →