Nvidia launches new security platform to keep AI agents under control following multiple breaches
Table of Contents
Nvidia has introduced the Open Agent Safety Platform, a new security platform designed to place stronger controls around autonomous AI agents. AI agents have become capable of doing much more than generating text, with some systems now able to use tools, access files, execute code, connect to external services, and complete tasks with limited human involvement.
That increased independence also creates new security risks, which have already been demonstrated by rogue OpenAI models. Nvidia says several AI labs have reported cases where agents moved outside the environments they were supposed to operate within and accessed systems they weren’t intended to reach. The company argues that relying on an AI agent to control its own behavior isn’t enough, particularly when agents are given more tools, more time, and broader access.
Nvidia introduces the Open Agent Safety Platform
Nvidia’s approach is therefore to put security controls around the agent rather than relying entirely on safeguards built into the AI model itself. A major part of the platform is Nvidia OpenShell, an open-source runtime released under the Apache 2.0 license. OpenShell runs autonomous agents inside isolated sandboxes and allows operators to define exactly what an agent is permitted to access. This can include restrictions covering files, network connections, processes, tools, and credentials.
Latest PC & Tech deals
- AMD Ryzenâ„¢ 9 9900X 12-Core - was $499 now $337
- iBUYPOWER Element Gaming PC - was $1,999 now $1,699
- LG 27GX700A-B 27-inch Ultragear- was $599 now $485
- Samsung SSD 9100 PRO 1TB - was $339 now $249
- HP OmniBook 5 16 inch - was $961 now $799
Prices correct as of September 21st, 2026.
These restrictions are enforced outside the agent itself. This is important because even if an agent behaves unexpectedly or tries something outside its assigned task, it doesn’t automatically have permission to access everything available on the system.
Nvidia compares the idea to the security model that eventually became important for web browsers. Instead of simply trusting every website to behave properly, browsers isolate pages and restrict what they can access. Nvidia wants a similar layer of protection around autonomous AI agents.
OpenShell is also designed around a zero-trust approach. An agent receives the permissions required to complete its task rather than unrestricted access to the host system. Nvidia’s documentation says the runtime can restrict filesystem access, prevent unauthorized network connections, protect credentials, and block privilege escalation.
For organizations that want another layer of security, Nvidia has introduced Sentry, which extends monitoring and enforcement into its BlueField hardware. This allows agent activity to be monitored separately from the system where the agent itself is operating.
The idea is to detect what Nvidia calls “drift,” where an agent begins taking actions that move away from its intended task or operating limits. This could happen because of unclear instructions, missing tools, software problems, or an agent repeatedly attempting different approaches to solve a difficult task.
Nvidia says its broader safety architecture is divided into three layers: the application containing the agent and its tools, the runtime responsible for monitoring and enforcing policies, and the underlying infrastructure running the workload. The company is also working with more than 100 industry partners around the new platform.