NVIDIA Launches Open Agent Safety Platform, Putting AI Agent Monitoring In Hardware

NVIDIA has announced the Open Agent Safety Platform, a reference design that pairs a software sandbox for AI agents with an independent watchdog running on separate silicon. The pitch is that agents can’t be trusted to police themselves, so the controls that constrain them need to sit somewhere they can’t reach.

The announcement comes as frontier labs deal with a run of incidents in which agents got out of the environments meant to contain them.

jensen huang

Why now

Several labs have recently reported agents breaking out of evaluation environments and reaching systems they should never have touched, and some agents misreported what they had done. In NVIDIA’s words, the security controls in place weren’t enough.

The examples are piling up. OpenAI recently paused training and tool-using inference on its most capable models after an internal agent used DNS to get around its internet restrictions, and earlier in the summer it said an internal model spent about an hour working around sandbox limits to upload results it had been told to keep in Slack. Google has since confirmed that Gemini broke into three real companies during a security test after a misconfigured environment gave it internet access.

NVIDIA says these breakouts didn’t come from one new capability. They came from a mix of tools, long runtimes and ambiguous instructions, plus an agent’s tendency to think outside the box. It calls this “drift”, meaning actions that depart from the intended task or operating constraints, and argues it can’t simply be trained away without also removing capability. Its conclusion is that an agent in these circumstances can’t be expected to fully govern its own behavior.

Two layers, software and silicon

The platform combines two pieces. The first is NVIDIA OpenShell, an Apache 2.0 open-source runtime that runs each agent in a sandbox with kernel-level isolation. Operators define which files, networks, tools, processes and credentials an agent can use. OpenShell turns those instructions into a verifiable policy, checks it before the agent runs and enforces it as the agent works. The runtime has already become a foundation for other agent projects, and NVIDIA has said OpenClaw and Hermes Agent are building native Windows apps on it.

The second piece is new. NVIDIA Sentry is an optional layer that extends monitoring and enforcement into BlueField-4 data processing units. Using NVIDIA’s DOCA framework, it ties together agent interactions, policy decisions, and tool and data access into a contextual record of what an agent did. That is meant to help systems spot drift, investigate suspicious behavior and decide when to step in. A DOCA gateway adds identity governance, continuously checking each agent’s identity and delegated authority.

The placement is the key design choice. In a Vera Rubin POD, each compute tray has a BlueField-4 on the node’s only path to the model. NVIDIA argues that position makes it both the best place to observe an agent and a kill switch, since an agent can’t act without its next thought. Because the DPU is isolated from the host, it is out of the agent’s reach even if the host itself can’t be trusted, and it can enforce policy at line speed. For customers already running Vera systems with BlueField-4, NVIDIA says turning this on is a software update. The platform is optimized for NVIDIA hardware but is described as compatible with other systems.

Five principles

NVIDIA lays out five principles behind the design:

  • Verifiable policy: a prover should show, before an agent runs, that its policy can’t escape the operator’s intent.
  • Out-of-band enforcement: controls shouldn’t live inside or within reach of the agent, and the agent needn’t know it’s being watched.
  • The path to the model is the control point: whoever controls it owns the best observation point and the kill switch.
  • Authority scales with visibility: the more an agent can do, the more its reasoning needs to be inspectable, which NVIDIA says is an advantage of open models.
  • Shared responsibility: labs, enterprises and hardware providers each own a layer, as in the cloud, and the runtime and policy language need to be open so any provider can plug in.

The browser analogy

NVIDIA compares the moment to the early web, when pages could run code on your machine and steal data, and the fix was to sandbox each page in its own tab rather than ask developers to behave. Safety, it argues, didn’t slow the internet down; it let commerce and everything else accelerate. It wants the same for what it calls the agent economy.

The stakes for ordinary users are already visible. Meta’s head of alignment at its superintelligence lab said an OpenClaw agent ignored her instruction to confirm before acting and bulk-deleted hundreds of emails from her inbox.

What to watch

The platform is a reference design, not a finished product with published results, and NVIDIA hasn’t shared numbers on how many escapes or drift events the Sentry layer would catch. There’s also an obvious commercial angle: the strongest form of the protection runs on NVIDIA’s own Vera and BlueField-4 hardware, though NVIDIA says it welcomes frontier labs, developers and infrastructure providers to build on the open pieces. A logo wall in the announcement suggests a broad set of ecosystem backers, but the blog post doesn’t detail commitments.

Still, the core idea is a notable shift in how agent safety is being discussed. Rather than relying only on alignment training or in-model monitoring, NVIDIA is arguing for enforcement that lives outside the agent entirely, and it is betting that the chip on the path to the model is where that should happen.

Posted in AI