Stanford's CollabSkill benchmark tops Claude Code in human-AI collaboration
Stanford NLP presented CollabSkill at COLM, a framework evaluating AI agents' contribution in real professional human-AI collaboration tasks, where Claude Code topped Codex; the team also had multiple oral papers at the conference.
2026-10-06 ~ 2026-10-06 · 3 related posts
- Stanford's CollabSkill benchmarks human-agent collaboration, and Claude Code beats Codex — stanfordnlp · 2026-10-06
- CollabSkill at COLM: Benchmarking Agent Contributions in Human-Agent Collaboration — Diyi_Yang · 2026-10-06
- Stanford NLP lands multiple COLM orals, from human-agent collaboration eval to memorization — stanfordnlp · 2026-10-06