Scale CEO touts assistant benchmark: Muse scores 9.3, beats Instinct 4-1 on real tasks
alexandr_wang · x · 2026-09-09
Scale CEO Alexandr Wang amplified a head-to-head on assistantbenchmark.com where the AI assistant Muse beat Instinct 4-1 (one tie), scoring 9.3 vs 8.6 across 16 equally-scaled real-world tasks.
- Online tasks: Muse scored 8, proactively offering four refundable hotel options when Chicago inventory was tight and asking whether to adjust budget or dates.
- Travel booking: Instinct took the win with a 10 — it logs into the user's own airline/hotel accounts for full supply, handles check-in and boarding passes; Muse relies on Duffel, so supply is thinner.
- Purchasing & recommendations: Muse shone with itemized totals before charging, virtual cards to hide real card numbers, and constraint-aware restaurant picks.
Note the source is Scale's own CEO promoting its benchmark and assistant, but the per-task breakdowns are concrete enough to serve as an agent-capability comparison reference.
More from Models
- Astra makes a weird dashboard mistake at just 44% context usage — eigenron · 2026-09-09
- NVIDIA open-sources gold-medal IMO system Nemotron with models, datasets and 200 new problems — kuchaev · 2026-09-09
- Claude asks heavy user for government ID to complete cyber verification — Bulky-Priority6824 · 2026-09-09
- Two years from o1-preview to superhuman math: RL scaling now cracks open research problems — jam3scampbell · 2026-09-09
- Meta's AI assistant outperforms Instinct and ChatGPT Work on a 5-task test, claims reviewer — garrytan · 2026-09-09
- Robot arm self-calibrates with 3 uncalibrated cameras, hits sub-0.2mm accuracy — burny_tech · 2026-09-09