Stanford Paper Introduces CollabSkill: Evaluating Human-Agent Collaboration
paraschopra · x · 2026-07-30
Stanford University and other institutions released the CollabSkill framework, designed to evaluate human-agent collaboration on real-world occupational tasks.
- Background: Existing AI benchmarks mostly focus on fully autonomous model performance, neglecting the human-agent collaboration paradigm, which better preserves human agency and generates economic value.
- Mechanism: The framework pairs real human workers with AI agents and employs a Bayesian skill rating system to disentangle and quantify the respective skill contributions of humans and AI.
- Key Findings: Based on data from 93 workers across 386 sessions, the study reveals that collaboration rankings diverge significantly from fully autonomous benchmarks (e.g., Claude Code ranks first in collaboration, whereas Codex leads in autonomous benchmarks). Furthermore, practical hands-on experience is the primary driver of collaboration skill.
Related event: Human Baselines Missing in AI Evaluations, Highlighting Human-AI Synergy(4 posts)→
More from coding & agent
- Debate Erupts Over OpenAI vs Anthropic Default Chain-of-Thought Retention in APIs — steipete · 2026-07-31
- Andrew Ng Uses Coding Agents to Turn Slides Interactive, Run LLMs in Browser — yuntiandeng · 2026-07-31
- Developer Successfully Boots Custom Robot Powered by Orin Nano Super — chrismatthieu · 2026-07-31
- Atelier: A VS Code Extension Bridging Specs and Kanban for AI Coding — narphorium · 2026-07-31
- Why Multi-Agent Coding Turns You Into Middle Management: Decomposition Is Key — Maxulis · 2026-07-31
- Building Agent Memory Systems with LangSmith and OpenWiki — BraceSproul · 2026-07-31