Assistant Benchmark goes live: 61 assistants scored across 15 real-use dimensions

Scobleizer · x · 2026-09-11

Assistant Benchmark v0.1 is live, scoring personal AI assistants after real use. First scorecard covers 61 assistants across 15 dimensions (online tasks, travel, email, purchasing, memory, phone calls, multi-step, etc.), with 13 of 61 tested as of Sep 10, 2026. Muse leads at 9.1 (7s responses), Instinct 8.3, szn 7.6; Grok Bot scores 7.3, while Catch, tinyNature and Town trail. Detailed test notes capture real failures, e.g. Shuffle couldn't reach Notion, Calendar or Slack and couldn't place calls. Feedback on methodology is being solicited.

Original post →

More from Models

Models channel →