Meta details Muse agent safety: isolated cells, no real credentials, and a Sentinel agents can't override
alexandr_wang · x · 2026-09-13
Meta Superintelligence Labs published a deep dive on how it built safety into Muse, its personal agent; Alexandr Wang says security was the longest pole before release.
- Model side: trained for zero-shot CLI/skills tool calling, long context, long-trajectory instruction following with inherent prompt-injection awareness, and multi-agent coordination.
- System side assumes the agent is always under attack: the harness runs in an isolated cell with no access to real credentials, and every external interaction flows through a Sentinel the agent cannot override.
- Hardening drew on dogfooding, agentic red teaming, and a private bug bounty; Meta is opening the Muse bounty program. Jason Lemkin praised the Secrets Store as fixing vibe coding's plaintext-credentials problem.
More from coding & agent
- mitsuhiko: AI Agents Can Fix Even a Broken Linux Install via Recovery Console — mitsuhiko · 2026-09-14
- Codex fixed his webcam: using coding agents to tame Linux — richardreis · 2026-09-14
- Dev: With Agents, You Can Fix Anything on Linux With a Prompt — steipete · 2026-09-14
- Marmel 0.9.0: autonomous coding agent tuned to run local models like gemma 4 12b — Naiw80 · 2026-09-14
- Which coding agent is everyone actually using? bough reads agent session histories — VoidEqualZero · 2026-09-14
- Dev builds agent harness for RTS games, optimizing info density and decisions per minute — coding_is_tedious · 2026-09-14