LangChain uses 'human touches per PR' as a key metric to evaluate coding agents
LangChain · x · 2026-10-01
Software engineering is shifting toward 'factory engineering,' where throughput matters and 'human touches per PR' serves as an efficiency metric to drive down over time.
LangChain's sydneyrunkle notes that scoring PRs is hard but tracking human touches per PR is easy, and her team uses this metric to evaluate coding agents. She cautions that on complex tasks, more touches isn't necessarily bad even when the agent performs well.
More from coding & agent
- rabbit OS3 ships 41 updates in 7 days, tops 100B tokens since launch — jesselyu · 2026-10-01
- Developer argues 100-1000 tps LLM speed is pointless unless rewriting legacy code to Rust in one pass — ssh4net · 2026-10-01
- BAAI's AREX-2 Trains Self-Improving Agents, Hits 92.2 on GAIA and 81.8 on MLE-bench Lite — BAAI · 2026-10-01
- Box^2-Bench Shows Frontier Models Struggle to Reject Unreliable Workflow Guidance — Minghan Wang · 2026-10-01
- SkillSeek: plain BM25 matches LLM-mediated agent skill retrieval at half the cost — StevensAGI · 2026-10-01
- Meta-Skill: Frozen-Weight Builder Models Learn Better Agent Harnesses — apodex · 2026-10-01