Assistant Benchmark goes live: 61 assistants scored across 15 real-use dimensions
Scobleizer · x · 2026-09-11
Assistant Benchmark v0.1 is live, scoring personal AI assistants after real use. First scorecard covers 61 assistants across 15 dimensions (online tasks, travel, email, purchasing, memory, phone calls, multi-step, etc.), with 13 of 61 tested as of Sep 10, 2026. Muse leads at 9.1 (7s responses), Instinct 8.3, szn 7.6; Grok Bot scores 7.3, while Catch, tinyNature and Town trail. Detailed test notes capture real failures, e.g. Shuffle couldn't reach Notion, Calendar or Slack and couldn't place calls. Feedback on methodology is being solicited.
More from Models
- ValsAI: Astra first AI to reach Minecraft Nether Fortress in long-horizon eval — scaling01 · 2026-09-11
- Dev: You Can Tell Who Has a Real RL Pipeline Just From Model Outputs — teortaxesTex · 2026-09-11
- DeepSeek Flash impresses: non-sycophantic, argumentative, and blazing fast — oran_ge · 2026-09-11
- Gemini glitches into endlessly spamming the word 'shame' — tugkanintassagi · 2026-09-11
- Early take: DeepSeek V4.1 Solo beats Agent Teams and GLM 5.3 Flash on quality and cost — teortaxesTex · 2026-09-11
- What counts as an 'exchange'? Anthropic's 865K/day metric questioned — teortaxesTex · 2026-09-11