Scale CEO touts assistant benchmark: Muse scores 9.3, beats Instinct 4-1 on real tasks

alexandr_wang · x · 2026-09-09

Scale CEO Alexandr Wang amplified a head-to-head on assistantbenchmark.com where the AI assistant Muse beat Instinct 4-1 (one tie), scoring 9.3 vs 8.6 across 16 equally-scaled real-world tasks.

Note the source is Scale's own CEO promoting its benchmark and assistant, but the per-task breakdowns are concrete enough to serve as an agent-capability comparison reference.

Original post →

More from Models

Models channel →