Stanford's CollabSkill benchmarks human-agent collaboration, and Claude Code beats Codex

stanfordnlp · x · 2026-10-06

Stanford NLP presents CollabSkill at COLM's first oral session, a framework for evaluating human-agent collaboration on real occupational tasks:

The team hopes it spurs systematic evaluation of human-agent collaboration beyond fully-autonomous benchmarks.

Related event: Stanford's CollabSkill Benchmark Puts Claude Code on Top for Human-AI Collaboration(2 posts)→

Original post →

More from coding & agent

coding & agent channel →