CollabSkill at COLM: Benchmarking Agent Contributions in Human-Agent Collaboration
Diyi_Yang · x · 2026-10-06
The CollabSkill paper was accepted to COLM (review scores 8/7/9/8) with an oral spotlight, presenting a framework for evaluating agent capabilities and contributions in human-agent collaboration, with extensive insights on AI literacy on the human side.
The authors released the collected trajectories and the official CollabSkill rating implementation. Fun detail: the dataset was uploaded on 8/1 but only just linked to the project page—yet it already has 1700+ downloads, hinting at how many bots/auto research agents are scraping Hugging Face these days.
Related event: Stanford's CollabSkill benchmark tops Claude Code in human-AI collaboration(3 posts)→
More from coding & agent
- Creator hands repetitive workflow to Codex and GPT-6 Astra, keeps creative calls — socialwithaayan · 2026-10-06
- Monica CRM MCP Server wraps REST API with 21 natural-language tools — modelcontextprotocol · 2026-10-06
- React Compiler core member's go-to prompt: make AI restate your goals before it works — sujingshen · 2026-10-06
- dotey explains Codex Project vs Claude Projects: forum boards vs Slack channels — dotey · 2026-10-06
- A Planted 'P.S.' Fooled Jev, TypeSafe's New Decision Model — a Simple Rule Caught It — Internal-Lie-5197 · 2026-10-06
- Claude Code's New Dreaded Message: 'Compacted While Idle, Before the Prompt Cache Expired' — dSebastien · 2026-10-06