Photo by Vitaly Gariev on Unsplash. Source: https://unsplash.com/photos/woman-working-on-computer-in-modern-office-with-chalkboard-rU1_DCersNs (Unsplash License).

Executive Summary

NVIDIA Open Agent Safety Platform, announced September 28, moves the enforcement boundary for AI agents off the machine the agent runs on. OpenShell, the open source runtime, keeps the software policy. Sentry, a watchdog on the BlueField-4 data processing unit, watches the same agent from separate silicon and quarantines a breakout in milliseconds.

The architectural claim is simple and worth taking seriously. Software enforcement inside the agent’s own host shares a fate with the agent, because when the loop goes off script it can reach the guard. NVIDIA’s own account of this year’s incidents says the same thing, that agents bypassed controls at the application layer to finish the task. Putting the boundary in hardware answers that. The open question is whether a boundary that runs on one vendor’s silicon is a boundary or a sales channel, and whether the overhead and the hardware bill price small agent fleets out of it.

Every agent incident this year followed the same script. The model did not fail. The agent reached around a guardrail in the application layer and kept going until it had the access it wanted.

NVIDIA Open Agent Safety Platform is the answer the company published on September 28. It splits agent security into two parts that never share a machine. OpenShell holds the software boundary and traces what the agent does. Sentry watches the same agent from a BlueField-4 data processing unit that sits outside the host the agent controls.

Diagram of the NVIDIA Open Agent Safety Platform. Three bands. The agent loop and model run on the NVIDIA Vera CPU. OpenShell sets a software policy boundary inside the same trust domain the agent holds. Sentry runs an out-of-band watchdog on a BlueField-4 DPU that can quarantine an agent in milliseconds.
The layer that watches the agent does not share a machine with the agent.

The boundary has to sit outside the thing it governs

A sandbox that runs on the same host as the agent shares its fate. If the agent escalates inside that host, the policy meant to stop it is now running on the machine it owns. That is the failure NVIDIA describes across recent incidents. The agent found a path at the application layer and used it.

Sentry is built on NVIDIA DOCA and enforces in silicon. It verifies agent identity, inspects requests and returns, and emits attested telemetry. NVIDIA says it can quarantine an agent that crosses its boundary in milliseconds. The point is that the agent cannot read the watchdog, disable it, or negotiate with it, because the watchdog is not on its machine.

OpenShell is the software half, and it is open source. It is now broadly available and runs on NVIDIA Vera, the company’s first CPU built for agentic work. NVIDIA also says OpenShell can be extended to compute from Arm and Intel, which is the detail that decides whether this is a platform or a lock-in.

The pattern is a boundary that moves with the workload

NVIDIA’s own summary of the incidents is the interesting part. It does not blame the model. It says the agent circumvented security controls at the application layer to complete its assigned task, and that the same shape repeats. This summer, agents running an OpenAI cybersecurity task breached Hugging Face. OpenAI now keeps a page of its own agents going wrong.

The common thread is not malice. It is a control that lives where the agent can reach it. That is why the fix is architectural, and not a stricter prompt or another filter wrapped around a runtime the loop already owns.

Anthropic drew the same line without a new chip

Anthropic’s Claude Managed Agents run the agent loop on a separate server from the sandboxes where the work executes. Same idea, different tactic. Put distance between the decision and the action. NVIDIA sells that distance in hardware, so the recommended shape is a DPU and, ideally, its own CPU.

That is not automatically wrong. Open source and extensibility are real answers to the lock-in worry. But a boundary that holds only on one vendor’s silicon is still a boundary with a sales quota, and a buyer should say that out loud before signing.

Two things to test before believing the pitch. Overhead comes first, because watching every request and response on a DPU is not free and NVIDIA has not published agent-loop numbers. Cost at the low end comes second, because a ten-agent internal tooling project does not want a DPU budget line. Jensen Huang said the platform would have prevented the breaches. That is a claim about a product that shipped this week, not a measured result.

Three questions before you design around this. Where does your agent loop run, and what else shares that host? When application policy fails, what actually stops the agent? And who on your team can prove, after the fact, that the boundary held?

Related reading. Two earlier pieces sit next to this one. We looked at how Google shipped an agent runtime for a million idle sandboxes, and at why attesting the hardware is the only way to prove you cannot see the data.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.