CollabSkill at COLM: Benchmarking Agent Contributions in Human-Agent Collaboration

Diyi_Yang · x · 2026-10-06

The CollabSkill paper was accepted to COLM (review scores 8/7/9/8) with an oral spotlight, presenting a framework for evaluating agent capabilities and contributions in human-agent collaboration, with extensive insights on AI literacy on the human side.

The authors released the collected trajectories and the official CollabSkill rating implementation. Fun detail: the dataset was uploaded on 8/1 but only just linked to the project page—yet it already has 1700+ downloads, hinting at how many bots/auto research agents are scraping Hugging Face these days.

Related event: Stanford's CollabSkill benchmark tops Claude Code in human-AI collaboration(3 posts)→

Original post →

More from coding & agent

coding & agent channel →