Gradescope cofounder benchmarks 5 multiplayer AI platforms on 18 weighted criteria
sergeykarayev · x · 2026-09-25
Sergey Karayev, cofounder of Gradescope, shares how his team evaluated multiplayer AI platforms using lessons from grading at scale: create descriptive rubric items, apply multiple items per evaluation, and crucially, adjust item weights on the fly.
The team defined 18 criteria (e.g. "in a shared conversation, the agent only uses connectors everyone present has access to") and evaluated Claude Tag, Dust, QM, Superconductor, and Viktor on each, recording videos as evidence. Results are presented in an interactive table where you can reweight criteria to get a score matching your team's needs.
More from Apps
- HeyGen survey: 53.9% of small business owners skip video over unprofessional look — Med1_Ai · 2026-09-25
- Open Notebook, an open-source NotebookLM alternative, hits 39.5k GitHub stars — adnan_hashmi · 2026-09-25
- Descript launches MCP to edit video from inside Claude, ChatGPT, and Codex — descript · 2026-09-25
- How one user trusts AI agents: refunds, price-triggered buys, and party planning — marilynika · 2026-09-25
- 10-minute workaround to register for Muse via cloud browser, card verification — TheMoonMidas · 2026-09-25
- Anyx quietly reads your WhatsApp groups and sends you a private daily AI summary — Scobleizer · 2026-09-25