ByteDance Seed's HarnessDev scores agents on harnesses they build, not tasks they finish
omarsar0 · x · 2026-09-02
ByteDance Seed's HarnessDev paper proposes scoring self-evolving agents on the execution harness they build rather than completed tasks.
Method:
- Agent starts from a weak runnable seed plus a few cases and builds a full execution system
- A second stage iterates the harness on downstream feedback
- Both stages scored on capability and execution-token cost
Experiments: 6 creator LLMs, 4 domains, 2,207 held-out instances.
Findings: generated harnesses trail mature human engineering on code and research, but match or beat it on writing and ML experimentation — a sharp domain split.
More from coding & agent
- Tigerless Labs' auto-gtm drafts X/Reddit posts from your PRs, never auto-posts — Aiden_Tech_Ai · 2026-09-03
- How to measure agent quality beyond task success: dev seeks real-world eval metrics — serpratik · 2026-09-03
- Giving your Grok bot an email and a credit card: webhook routines that buy things on Amazon — jeff_weinstein · 2026-09-03
- Dev builds full multi-agent stack on Nostr relay with encrypted A2A comms and unified MCP — RileyRalmuto · 2026-09-03
- Study of 8,351 Claude Code plugins finds 74% of docs commits are runtime instructions — rohanpaul_ai · 2026-09-03
- Codex hooks hide exit codes — stalegreen rewrites verification commands and blocks 26% of stale green claims — SmiLePLSSS · 2026-09-03