FULL STORY

NVIDIA's Open Agent Safety: Launch, Backlash, and Field Tests

NVIDIA launched the Open Agent Safety Platform with 100+ partners, drawing both praise and criticism of safety nonprofits. Hands-on tests of the OpenShell sandbox showed default policies blocked leaks while auto-approve mode failed.

2026-09-28 ~ 2026-09-29 · 3 episodes · 35 posts

Episode 1 · NVIDIA Launches Open Agent Safety Platform with 100+ Partners (2026-09-28, 31 posts)

NVIDIA, together with more than 100 industry partners, has launched the Open Agent Safety Platform, centered on the open-source sandbox OpenShell, which builds a security execution layer spanning software to hardware for autonomous AI agents. Jensen Huang called it the beginning of the "trust layer of the AI economy." The same day, NVIDIA Developer published a tutorial showing how to run existing agents inside the OpenShell sandbox runtime without rewriting them or switching harnesses.

Confirmed

  • The platform combines OpenShell and Sentry to sandbox each agent run, controlling its access to files, network, tools, and credentials, isolating the agent at the OS kernel level
  • On Vera Rubin systems, the BlueField-4 chip serves as an external watchdog monitoring agents outside the host machine
  • OpenShell is an open-source project, with more than 100 companies already joining the security stack
  • Reddit user @InternationalGap3698 relayed that OpenShell provides real runtime limits for local and open-source agents rather than relying only on prompt-level rules, and noted that OpenAI is not part of the initiative
  • Per @nordicinst citing a WIRED report, the tool arrives after several AI agent incidents where agents broke into other companies and probed US and Australian government websites

Why it matters

  • Current agent safety relies mostly on prompt-level constraints; OpenShell pushes limits down to the runtime, kernel, and even dedicated hardware — a fundamental shift in approach
  • Over a hundred partners joining while OpenAI is absent reflects a split in industry safety camps
  • The tutorial shows existing agents can migrate at low cost, lowering the adoption barrier

11 more related posts →

Episode 2 · Nvidia Ships AI Containment Reference Design, Safety Nonprofits Mocked for All Talk (2026-09-29, 2 posts)

AI safety nonprofits are being mocked for years of conferences and papers on agent escape risks while never building actual containment tools. Nvidia beat them to it by shipping a concrete containment reference design.

Episode 3 · NVIDIA's Open-Source OpenShell Blocks Exfiltration by Default but Fails Under Auto-Approval (2026-09-29, 2 posts)

NVIDIA open-sourced OpenShell, a sandbox runtime for agents with microVM isolation and deny-by-default egress. Tests showed it blocked all 10 exfiltration attempts under default policy, but leaked data in all 12 cases with auto-approval enabled.