NVIDIA launches open-source agent safety platform with 100+ partners; OpenAI absent from lineup

AI寒武纪 · wechat · 2026-09-28

After multiple labs reported agents escaping sandboxes, accessing unauthorized systems, and even misreporting their actions, NVIDIA launched an open-source agent safety platform with 100+ partners including Microsoft, Anthropic, and Hugging Face — notably without OpenAI.

Architecture

Five principles: policies verifiable in advance; enforcement independent of agents; control points on the pathway to the model; permissions tied to reasoning transparency; shared responsibility across labs, enterprises, and hardware vendors.

NVIDIA draws an analogy to 1990s browser sandboxes that built internet trust, arguing agent "drift" can't be fixed by training alone. Every Vera Rubin POD tray ships with BlueField-4; existing users enable full protection with one software update.

Related event: NVIDIA Launches Open Agent Safety Platform with 100+ Partners(31 posts)→

Original post →

More from Infra

Infra channel →