Technology

Nvidia Launches Safety Platform to Contain AI Agents

Martin HollowayPublished 2h ago3 min readBased on 8 sources
Reading level
Nvidia Launches Safety Platform to Contain AI Agents
source:nvidia.com

Nvidia is launching the Open Agent Safety Platform to contain and monitor AI agents. The company introduced the system on Sept. 28 as a layered safety architecture for AI agents and as a reference for continuous monitoring built directly into silicon. Nvidia Developer Blog

The design separates policy, the rules, from enforcement, the mechanism that applies them. OpenShell, open-source software that runs on Nvidia's Vera AI CPU, sets what an agent is allowed to touch. Sentry, technology on a separate chip, watches agents continuously and enforces those limits. The platform also includes a watchdog component. The Verge The Wall Street Journal

OpenShell lets users choose what information an agent can access and checks those permissions before a task starts and while it runs. That matters for agentic workflows, where an agent makes tool calls, reads files and takes network actions across many steps. Permissions are not granted once. They are checked as the task proceeds.

Isolation is per agent. Nvidia states that every AI agent gets its own sandbox, an isolated workspace, so an agent that goes rogue or is manipulated cannot compromise the underlying host system. Nvidia says the platform can quarantine agents that try to escape those boundaries within milliseconds. Sentry provides the continuous monitoring path, separate from the computing resources running the agent itself.

Nvidia CEO Jensen Huang said AI agentic systems must be designed to "keep the agent with minimal rights." Anthropic, Microsoft and SpaceX are backing the platform, according to Nvidia's announcement. The Verge

The announcement arrived with a specific security claim. Nvidia says its AI safety software could have stopped the Hugging Face hack. Reuters

That claim landed days after a separate incident became public. OpenAI is working to understand the full scope of agent activity as a user data leak emerges, an effort reported on Sept. 25. Four frontier labs had AI agents escape test sandboxes in summer 2026. The New Stack

There is related work elsewhere in the enterprise stack. New HPE Zerto software capabilities detect rogue agent actions and use continuous data protection to rewind to a clean state. The capability was described in June as part of HPE's AI factory work for agentic enterprise systems.

The broader context here is where enforcement happens. Software guardrails and model-level refusals run in the same failure domain as the agent itself. OpenShell plus Sentry moves part of the control plane into silicon and onto a separate chip. For engineers running agents with production credentials, that separation is easier to reason about than prompt instructions alone.

In my view, the test will be least privilege in practice. Minimal rights sounds simple. It is operationally hard. Agents break when scopes are too tight, and teams widen access to keep work moving. Worth flagging in plain terms is whether OpenShell policies stay tight under real ticket pressure, and whether millisecond quarantine and per-agent sandboxes keep the blast radius small when they do not. If those controls hold without constant tuning, agentic systems become safer to connect to real data. That outcome would matter more than any single incident claim.