Stanford's CollabSkill benchmark tops Claude Code in human-AI collaboration

Stanford NLP presented CollabSkill at COLM, a framework evaluating AI agents' contribution in real professional human-AI collaboration tasks, where Claude Code topped Codex; the team also had multiple oral papers at the conference.

2026-10-06 ~ 2026-10-06 · 3 related posts