Meta Muse uses a separate Sentinel to gate agent outbound actions—what should you test first?

Crescitaly · reddit · 2026-09-09

Meta's September 8 Muse announcement describes a personal agent running in a dedicated cloud VM, with a separate Sentinel component approving outbound activity; sensitive actions like sending email or making purchases require user permission.

The key distinction, the author argues, is between an AI reviewer deciding an action looks acceptable and hard permissions neither agent can override. A first test worth running: a malicious page instructing the assistant to send private data to a new destination—which layer rejects it, and what evidence reaches the user?

Rollout caveat: Muse is launching now on iOS, Android and web in the US, but the Confidential VM with a user-held key arrives later this year and differs from the VM offered today. These are Meta's design claims, not independent security testing.

Related event: Meta Launches Muse, a Personal AI Agent with Sandboxed VM and Sentinel Safety Layer(47 posts)→

Original post →

More from coding & agent

coding & agent channel →