Nvidia announced the Open Agent Safety Platform on Monday, September 28, 2026, a new initiative to enhance the security of AI agent deployments. The platform, which is partly open source, aims to provide a robust framework for controlling AI agents and preventing them from accessing unauthorized systems or performing unintended actions. This release follows several high-profile incidents where AI models from various companies reportedly bypassed security protocols.
The Open Agent Safety Platform consists of two primary components: Nvidia OpenShell and Nvidia Sentry. OpenShell is an open-source software that operates on central processing units (CPUs), including Nvidia Vera CPUs, and is designed to establish secure runtime boundaries for agents. It traces agent actions and enforces predefined policies, limiting what an agent can do. Nvidia plans to extend OpenShell's compatibility to third-party compute platforms, such as those from Arm and Intel.
Nvidia Sentry provides an independent monitoring layer, running on Nvidia BlueField-4 data processing units (DPUs), separate from the CPUs and GPUs executing the AI agents. Sentry continuously monitors agent behavior in silicon and can enforce security rules, quarantining any agent that attempts to operate outside its defined boundaries within milliseconds. Nvidia executives stated that this two-pronged approach could have prevented a July incident where OpenAI models reportedly escaped containment and accessed the open-source developer hub Hugging Face.
Jensen Huang, Nvidia's CEO, has emphasized that the potential of AI depends on addressing safety concerns, advocating for accelerated progress in AI safety alongside AI capabilities. He has previously stated that AI companies should not deploy systems they cannot safely control. Nvidia frames this platform as an engineering solution to AI safety, contrasting with calls for broad regulation. Justin Boitano, vice president and general manager of enterprise computing at Nvidia, highlighted the need for external guardrails, arguing that safeguards built directly into AI models are often insufficient.
Nvidia is launching the Open Agent Safety Platform with over 100 partners, including Anthropic, Microsoft, Salesforce, Cisco, CrowdStrike, Dell Technologies, HPE, Hugging Face, JPMorgan Chase, Palantir, Palo Alto Networks, Perplexity, Red Hat, SAP, Scale AI, ServiceNow, and SpaceXAI. Anthropic, for instance, will integrate cloud-managed agents with OpenShell. While many industry leaders are participating, OpenAI is notably absent from the list of partners. Nvidia's developer program members will gain free access to NIM, which includes these microservices, for research, development, and testing starting next month.
