ByteDance Seed's HarnessDev scores agents on harnesses they build, not tasks they finish

omarsar0 · x · 2026-09-02

ByteDance Seed's HarnessDev paper proposes scoring self-evolving agents on the execution harness they build rather than completed tasks.

Method:

Experiments: 6 creator LLMs, 4 domains, 2,207 held-out instances.

Findings: generated harnesses trail mature human engineering on code and research, but match or beat it on writing and ML experimentation — a sharp domain split.

Related event: ByteDance Seed Introduces HarnessDev: Agents Build and Evolve Their Own Harnesses(2 posts)→

Original post →

More from coding & agent

coding & agent channel →