Nvidia Puts AI Agents Under Separate Watch With New Safety Platform

Nvidia chief executive Jensen Huang introduced the Nvidia Open Agent Safety Platform on September 28, 2026, a toolkit of software and hardware that adds separate security layers around AI agents TechCrunch.
The platform has two parts. OpenShell is open-source software that sets what an agent can access while it works. Sentry is a separate monitoring system that watches that work and steps in when an agent moves outside its defined limits.
Sentry does not run on the same CPU or GPU as the agent. It runs on Nvidia's BlueField-4 data processing units, or DPUs, small processors built to handle data movement and infrastructure jobs. Because enforcement lives on a different chip and data path, the monitoring and quarantine logic keeps running even if the agent's own environment is compromised.
Nvidia says Sentry watches behavior continuously and can quarantine an agent that crosses its boundaries within milliseconds. AI agents can chain API calls, file access, and code execution in seconds, so response speed decides whether a rule violation stays contained or becomes a breach.
OpenShell was announced in March, ahead of the September platform launch. Nvidia also released NemoClaw in March, an enterprise-grade agent platform based on OpenClaw with built-in security. OpenClaw is an operating system for agents created by Peter Steinberger. Huang said work on the safety platform began a year ago, after the introduction of OpenClaw.
Nvidia states that autonomous AI agents rely on secure infrastructure layers, including sandboxes, or isolated work areas, identity controls, and policy engines, to manage tool access and protect sensitive data Nvidia. OpenShell covers the control side for access. Sentry covers the monitoring side. The setup follows a zero-trust idea, meaning no agent is trusted by default, which fits agents that hold broad tool permissions and persistent state.
Huang said in a CNBC interview that the platform would have prevented earlier AI agent breaches. Nvidia said its new safety software could have stopped the Hugging Face hack Reuters. Nvidia described the platform on September 28, 2026 as a security system intended to stop AI agents from going rogue.
Companies supporting the effort include Anthropic, Arm, Microsoft, Oracle, and SpaceX. OpenAI is not listed as a participating company. Containment only works if model providers, infrastructure vendors, and cloud operators agree on common formats for policy expression, attestation, and telemetry.
The broader context here is enforcement architecture versus regulation. Huang said on September 15 that the AI industry does not need any new AI security laws or regulations Bloomberg. In my view, that position is best read as an argument for platform-level controls over new statutes, not as a claim that risk is low.
What is worth flagging for security teams is what this design would allow if it works as described. Out-of-band monitoring on DPUs would let policy survive even if the agent is compromised. It would allow least-privilege tool access, identity-bound actions, and sandbox egress rules to be checked for each action rather than once per session. That is the missing layer for letting agents write to production systems, customer data, and outside services. We have seen this pattern before, when separation moved from hypervisors to sidecars to confidential computing, and agent infrastructure now appears to be following the same path.


