As AI intelligents move from answering questions to operating files, invoking tools, and accessing enterprise systems, security issues also shift from "saying the wrong thing" to "doing the wrong thing." On September 28th, NVIDIA announced Open Agent Safety Platform, aiming to isolate, monitor, and handle the operations of these intelligents within the same open architecture. The announcement included an open-source runtime OpenShell, as well as a reference design Sentry for independent monitoring. The focus of the product is not to ensure that every answer given by the model is correct, but rather to limit the actions that the model can take even if it is misled.
This difference is not abstract. An agent with permission to read emails and create external links may interpret instructions in malicious emails as user requests; a development assistant capable of executing commands might bring untrusted scripts into the local environment when installing dependencies. Past security tools were mostly designed around human accounts and deterministic programs, but agents can plan their steps autonomously and dynamically select tools on their own. Simply reminding users at the input stage to “not follow malicious instructions” is insufficient to replace the hard constraints in the execution environment.
OpenShell Put control into runtime.
NVIDIA describes OpenShell as an open-source runtime available for developers to use, which is used to manage the resources that agents can access and the actions they can perform. The approach is to provide agents with a restricted workspace and continuously apply policies when interacting with external systems. Checking permissions as files, network requests, and tool calls occur, rather than only after reviewing chat records, is closer to the existing principle of least privilege in enterprises.
The value of runtime constraints lies in their independence from whether the model “understands” a security prompt. Suppose a sales assistant is authorized to access a list of customers but does not have the right to send that complete list externally; even if it generates an upload request, the execution layer should still prevent unauthorized connections. The closer security controls are to the actual actions, the greater the chances of defending against prompt injection, misplanned operations, or abnormal returns from tools. However, errors in policy design can still leave vulnerabilities; what the platform provides is the capability for execution control, not the automatic creation of permission rules for each company.
The announcement states that OpenShell can be controlled on NVIDIA Vera CPU and can also be extended to Arm and Intel platforms. It is important to distinguish between "can be extended" and "all chip combinations have been verified by customers." Open source provides developers with the space to review and modify it, but it also means that companies need to check the versions, dependencies, patches, and operating environments themselves. Installing a set of open-source components in a test environment is different from running them stably in multi-departmental operations; these are two separate stages.
NVIDIA also proposed a reference design using BlueField-4 and DPU to act as gatekeepers outside of the main operating environment for the agents. The idea of independent channels is to retain the ability to observe and isolate in case the main environment is compromised or exhibits abnormal behavior. The announcement mentions a millisecond-level isolation goal, but the reference design cannot be equated with large-scale customer systems having already achieved the same effect; actual latency will still be affected by network conditions, the number of policies, application architecture, and load. Enterprises should request repeatable test data when making purchases, rather than treating the design goals as a commitment.
The security boundaries should be designed together with the business boundaries.
It is difficult to set the permissions for an agent all at once. If too few permissions are granted, it won't be able to complete even normal tasks; if too many are granted, a single misjudgment could lead to data leakage or incorrect transactions. In practice, tasks can be divided into four layers: reading, providing recommendations, preparing for execution, and final submission. The first two layers allow for more automation, while the latter two, which involve funds, sensitive data, or irreversible modifications, should retain approval and logging processes. For runtime systems like OpenShell, to truly be effective, they must work in conjunction with the organization's identity management system, data classification, and approval procedures.
Monitoring should not merely record the model's inputs and outputs. An agent may sequentially invoke search, database, and email tools; each individual request may seem reasonable on its own, but when combined, they can lead to unauthorized outcomes. Security teams need to see the complete call chain, the identity used at each step, the actions that were blocked, and whether retries were attempted after exceptions occurred. Otherwise, incident reviews will still get stuck at the question of "why did the model think that way," without a clear understanding of what the model actually did.
NVIDIA's business motivations in this area are clear: the more agent applications there are, the more enterprises need stable, isolated, and secure infrastructure. However, the effectiveness of security cannot be proven solely by vendor press releases. Customers should measure false positive rates, false negative rates, recovery times, and operational costs under real permissions, with real tools, and using adversarial samples. It is especially important to test scenarios where the security measures themselves fail. Open source and hardware isolation can increase options, but they do not replace these validation processes.
This release marks the beginning of a shift in infrastructure competition for AI, where "what is allowed to be done" is now considered as important as "how fast it can be done." OpenShell has been made available as an open component, while Sentry remains a reference design; their levels of maturity differ. For companies preparing to deploy agents, the most useful conclusion is not to claim that risks have been resolved, but rather to define clear permission boundaries within a system that is executable, testable, and accountable, before deciding how far the agents can go.
The deployment pace should also match the level of risk. Enterprises can first allow agents to process public information in an isolated environment, record their reactions when encountering malicious websites, phishing emails, and feedback from faulty tools, and then gradually integrate them into internal systems. With each additional permission granted, it is necessary to confirm the authorized person, the method for revoking that permission, and the recovery process in the event of an incident. Security platforms can provide control points, but whether these control points are effectively utilized depends on the operating procedures jointly established by business leaders and the security team.












