NVIDIA said on Sept. 28, 2026, that it is releasing Open Agent Safety Platform, an open software platform and reference system design for securing AI agents from testing through deployment. The company is not shipping a new model or assistant. It is trying to define a control layer around agents: where they can act, what they can touch, and which system can still stop them if they drift outside policy.

That matters because agent safety has moved from a theoretical concern to an operational one. Once an agent can call tools, query services, or access internal data, a model that is “well behaved” in a prompt test can still act in risky ways at runtime. NVIDIA’s announcement is built around the argument that model-level safeguards are not enough on their own. CNBC’s report on the launch described the product the same way: a new platform for setting safeguards that prevent agents from breaking out of containment, with NVIDIA framing the problem as an engineering one rather than a purely policy one.

What NVIDIA says is inside the platform

According to NVIDIA’s release, the platform has two main pieces. OpenShell is open-source runtime software that traces agent actions and enforces policy on NVIDIA Vera CPUs. NVIDIA says it can also be extended to third-party compute platforms, including Arm and Intel systems. In practical terms, that makes OpenShell the software boundary: it is meant to define what an agent is allowed to do while the task is running, not just what it was instructed to do at prompt time.

The second piece is Sentry, a reference system design built around NVIDIA BlueField-4 DPUs. NVIDIA says Sentry runs as an out-of-band watchdog that continuously monitors agent behavior and can quarantine an agent in milliseconds if it tries to move outside its boundary. The distinction matters. If the runtime boundary is breached, the watchdog is intended to sit in a separate trust domain and still be able to enforce policy. That is the architectural bet behind the launch: safety controls should not depend only on the same software stack the agent is trying to use.

NVIDIA also says the platform uses DOCA software to inspect requests and responses, verify agent identity, provide telemetry, and enforce zero-trust access policies across data, tools, APIs and services. That points to a broader design goal: not only blocking obviously malicious actions, but making agent activity auditable enough for enterprise security teams to review permission changes and access patterns after the fact.

Why enterprise and infrastructure teams should care

For enterprises, the immediate consequence is not abstract safety branding. It is the possibility of putting a containment layer between an agent and the systems it can reach. That is especially relevant for teams that want agents to work across software development, customer support, business workflows or robotics, but cannot afford broad, permanent access to credentials and internal services.

NVIDIA says more than 100 organizations are working with the platform technologies, including Anthropic, Cisco, Microsoft, Oracle, Salesforce, SAP, Scale AI and Red Hat. Those partnerships suggest the company is trying to make the platform an ecosystem layer rather than a standalone product. CNBC reported that Nvidia is also working with Anthropic to connect cloud-managed agents with OpenShell. NVIDIA’s own release gives examples such as Salesforce integrating OpenShell with Slack approvals and SAP embedding it with Joule Studio runtime. Those integrations matter because they show how the controls could fit into existing approval and audit workflows instead of requiring a separate console for every agent.

That is also where the practical decision begins for buyers and builders. If an organization is already deploying agents in sensitive environments, the key question is no longer whether the model is capable, but whether the runtime boundary is enforceable outside the model and outside the app layer. NVIDIA is explicitly arguing for that shift. For security teams, the value proposition is containment plus auditability. For platform teams, it is a way to standardize permission handling across different agents and workloads.

What remains unproven

The launch is still a vendor announcement, not an independent validation of safety or performance. In the supplied material, NVIDIA does not provide third-party benchmarks, comparative testing, pricing, or a public study showing that the platform prevents real-world incidents better than existing controls. The claims about minimal overhead and millisecond quarantine come from NVIDIA’s own release. As with most reference designs, the outcome will depend on how partners implement the stack, how consistently policies are written, and whether organizations actually place their trust boundary outside the agent itself.

That limitation is not trivial. Open source can speed adoption, but it can also push implementation complexity onto customers and partners. A control layer only works if it is integrated into the systems that grant access, log activity, and escalate exceptions to humans when needed. In other words, the platform may be strongest where an organization already has mature identity, policy and audit processes, and weaker where teams are still experimenting with loosely governed agents.

The launch therefore says as much about the direction of AI infrastructure as it does about safety. NVIDIA is positioning agent security as a stack problem: hardware enforcement, runtime policy, observability and human oversight. If that framing gains traction, the next competitive line in enterprise AI will not just be which agent is most capable, but which one can be contained, verified and governed when it acts on the company’s behalf.

If your agents can touch tools, APIs, or internal data, the key signal to watch is whether policy enforcement sits outside the model and sandbox, because that is the control boundary NVIDIA is trying to move.