Stanford's CollabSkill benchmarks human-agent collaboration, and Claude Code beats Codex
stanfordnlp · x · 2026-10-06
Stanford NLP presents CollabSkill at COLM's first oral session, a framework for evaluating human-agent collaboration on real occupational tasks:
- Pairs real workers with AI agents on tasks matched to their occupations, using a Bayesian skill-rating system to disentangle human vs. agent contributions
- Built on 1,500+ prompts from 386 working sessions by 93 human workers
- Agent-side rankings diverge sharply from fully-autonomous benchmarks (where Codex leads): Claude Code ranks first
- Human-side: practical experience is the primary driver of collaboration skill, and hands-on collaboration meaningfully shifts AI literacy
The team hopes it spurs systematic evaluation of human-agent collaboration beyond fully-autonomous benchmarks.
More from coding & agent
- Burkov's post-update workflow: fast tasks vs slow tasks now split across differently named model tiers — burkov · 2026-10-06
- PluginRSI: recursive harness improvement via reusable plugins beats whole-program search — Yaorui Shi · 2026-10-06
- ASCENT: online test-time training lets LLM agents self-distill verified experience into weights — Haodong Lu · 2026-10-06
- Open-source AI Engineering from Scratch curriculum builds every algorithm from raw math, agent-tutored in your terminal — ghumare64 · 2026-10-06
- Comparing 4 AI Animation Pipelines with Opus 5.5: Token Usage Turns Out Similar — chongdashu · 2026-10-06
- Cal.com Lands in Grok Build Plugin Marketplace With OAuth, No API Key Needed — tetsuoai · 2026-10-06