Gradescope cofounder benchmarks 5 multiplayer AI platforms on 18 weighted criteria

sergeykarayev · x · 2026-09-25

Sergey Karayev, cofounder of Gradescope, shares how his team evaluated multiplayer AI platforms using lessons from grading at scale: create descriptive rubric items, apply multiple items per evaluation, and crucially, adjust item weights on the fly.

The team defined 18 criteria (e.g. "in a shared conversation, the agent only uses connectors everyone present has access to") and evaluated Claude Tag, Dust, QM, Superconductor, and Viktor on each, recording videos as evidence. Results are presented in an interactive table where you can reweight criteria to get a score matching your team's needs.

Original post →

More from Apps

Apps channel →