Meta spent 7 months and 100+ agent evals polishing Muse after OpenClaw's reliability failures
Hesamation · x · 2026-10-12
According to Alexandr Wang, Meta has been building the Muse agent since February: the prototype took just 1-2 weeks, but the team then spent 7 months on reliability. A spreadsheet tracked 100+ specific agent behaviors, each with its own eval and a launch-blocking threshold — only Muse Spark 1.3 turned all 100 green, with under 200 people across model and product. The lesson from OpenClaw, which had a magic moment but broke constantly: reliability IS the product. Whether it holds up for millions of users is another question.
More from coding & agent
- Handy Claude Code trick: just reference @AGENTS.md from your CLAUDE.md — weswinder · 2026-10-12
- Custom church payment system: GPT security audit caught a critical flaw before it was too late — petergyang · 2026-10-12
- AI Agent Does CAD: KiCad MCP Reads Mounting Holes for a Robot Build — burhop · 2026-10-12
- Opus 5.5 Compared: Claude Teammates Push Back, Codex Quietly Drifts Behind Green Tests — Sauers_ · 2026-10-12
- AI Writes Code Faster — So Why Aren't We Shipping Faster? — kristiyanstoyanovAI · 2026-10-12
- Enterprise AI Rollouts: 5 of 300 Users, and What Clients Still Can't Get — Initial_Orange2985 · 2026-10-12