The Agent Security Stack Is Moving Below the Model
NVIDIA’s Open Agent Safety Platform exposes the deeper problem with autonomous AI: useful agents are becoming privileged software, and prompts are not a security boundary.

What changed
NVIDIA has launched the Open Agent Safety Platform, pairing the open-source OpenShell runtime with a Sentry reference design built around BlueField-4 DPUs. The headline sounds like another entry in the rapidly filling AI-security category. The architecture is more consequential than that. NVIDIA is arguing that the security boundary for autonomous agents cannot live primarily inside the model, prompt, or agent harness. It has to move beneath them. 1
OpenShell 0.1.0 puts an externally enforced runtime around an agent without requiring the agent itself to be rewritten. Its Gateway manages sandboxes and policy; a Supervisor sits outside each agent workload and inspects outbound requests; and the Sandbox applies kernel-level restrictions to files and processes while forcing network traffic through the supervisor. Policies can distinguish a read from a write against the same API. Real credentials can remain outside the agent workload and be substituted only for approved destinations and operations. 2
That distinction matters because the agent does not get to be the final authority on its own permissions. OpenShell can keep controls in force when an agent opens a shell, executes generated code, spawns child processes, delegates work to sub-agents, or asks for expanded access. NVIDIA has also added formal policy analysis intended to detect alternate paths through which apparently narrow permissions can combine into a prohibited action. 2
Sentry pushes the same idea farther down the stack. NVIDIA describes BlueField-4 DPUs as an independent, hardware-isolated enforcement and monitoring layer capable of observing and quarantining workloads outside the host execution environment. 1 IBM, meanwhile, is integrating agent identity, HashiCorp Vault, OpenShift, storage and infrastructure controls with OpenShell, an important sign that the design is already being treated as an enterprise control-plane problem rather than simply an NVIDIA developer feature. 3
Independent reporting confirms the launch and its security motivation. 4
The real problem: useful agents are becoming privileged software
The dirty secret of enterprise agents is that the industry has been trying to secure a new class of privileged software with controls inherited from chatbots.
That worked tolerably when an AI system mostly generated text for a human to inspect. It becomes much less convincing when an agent can authenticate to systems, query confidential data, modify repositories, invoke MCP servers, call business APIs, run arbitrary code, provision infrastructure, delegate work, and continue operating for hours or days.
Those capabilities are not edge cases. They are the point of agentic computing.
The more useful an agent becomes, the larger its authority surface becomes. Giving an agent access to a credential does not merely give it access to a secret; it may give it every operation that credential permits. Giving it a network path to an API may expose both read and write operations. Allowing generated code creates execution paths the original application developer never explicitly authored. Sub-agents create another problem: individually reasonable permissions can compose into a system-level capability nobody intended.
Traditional model guardrails attack a different layer of the problem. They try to influence what the model decides. Agent-harness controls can restrict which tools are presented. Application authorization can limit the user. All remain useful. None, by itself, establishes an independent security boundary around a probabilistic system that can reason about its own constraints.
That is the conceptual leap behind OpenShell.
Intelligence can remain probabilistic. Authority cannot.
Why the existing agent-security model is insufficient
Enterprise security normally assumes that important controls should not depend entirely on the workload being controlled. Operating systems do not ask applications whether memory isolation should apply. Network policy does not rely on a compromised server voluntarily declining an outbound connection. Secrets managers exist precisely because applications should not scatter durable credentials through their own execution environment.
Agents are forcing AI architecture to rediscover these principles.
NVIDIA's examples make the problem concrete. OpenShell can let an agent perform a GET against an API while blocking a POST even if the underlying credential itself has write permission. It can substitute the real credential outside the workload so the agent never receives the raw secret. It can reject a credential if an agent attempts to send it to an unapproved endpoint. And when an agent discovers it needs additional access, it can propose a policy change without being allowed to approve that change itself. 2
The last point may prove especially important. Autonomous agents will encounter situations designers did not anticipate. A system that simply fails whenever permissions are incomplete will not be useful enough. A system that lets the agent expand its own permissions is not a serious security architecture. The emerging answer is negotiated autonomy: the agent can identify what it needs, but a separate authority evaluates and grants the capability.
NVIDIA is even applying formal methods to this problem. Its policy prover analyzes the modeled permissions themselves rather than accepting the agent's explanation of what those permissions supposedly mean. NVIDIA says adversarial experiments included agents spending extended periods attempting to persuade an AI reviewer to grant permissions that could modify a protected repository; policy analysis was used to establish what the proposed permissions actually allowed. 2
That is a subtle but important design principle. The security system should evaluate capability, not intent.
A new control plane is forming around agents
The larger consequence is that enterprise agent architecture is beginning to separate intelligence from authority.
The model reasons. The agent plans. Tools execute. But an independent control plane increasingly determines which identities, credentials, services, files, processes and actions are actually available.
This is starting to look less like chatbot governance and more like a fusion of container security, zero trust, privileged-access management, API security and workload orchestration.
NVIDIA's own earlier secure-agent architecture points in this direction: managed workspaces, SSO, default-deny network access, centralized audit, protected credentials, runtime sandboxing and continuous policy verification. 2 IBM's same-day integrations reinforce it by tying agent identity and secrets management directly into the runtime boundary. 3
For CIOs, this suggests that "the agent platform" will not ultimately be one product. Enterprises are likely to need at least an intelligence layer, an agent/orchestration layer, and an independent authority layer. The last one may become strategically important because it can govern agents from multiple model and application vendors.
That also makes portability a first-order concern. If policy becomes inseparable from one model provider or one agent framework, enterprises will recreate the lock-in they are currently trying to avoid at the model layer.
The part NVIDIA is not emphasizing
NVIDIA's architecture is technically credible. It is also strategically convenient for NVIDIA.
If the industry's answer to agent risk is primarily better prompts and model-level guardrails, much of the control plane belongs to model and application vendors. If safe autonomy instead requires sandbox runtimes, protected credential paths, policy enforcement, trusted infrastructure and eventually hardware-isolated monitoring, the center of gravity moves toward infrastructure.
That expands NVIDIA's role from supplying the compute on which agents run to helping define the trust boundary around what agents are allowed to do.
Sentry is particularly revealing. A BlueField DPU that independently observes and constrains agent workloads makes the infrastructure itself part of the AI governance architecture. For highly regulated or high-consequence environments, that could be compelling. It could also create a new form of infrastructure dependency if enterprises allow policy semantics to become tightly coupled to a particular hardware stack.
CIOs should therefore separate the architectural principle from the vendor implementation.
The principle is strong: agent authority should be externally enforceable, least-privileged, auditable and resistant to manipulation by the agent itself.
Whether OpenShell and Sentry become the dominant implementation is a different question.
Frontier take
The agent-security battle is moving below the model.
For the last several years, much of AI safety has concentrated on making models behave: alignment, prompt controls, content filters, tool restrictions and increasingly elaborate instructions about what an AI should or should not do. Agentic systems expose the limit of that approach. Once software can act, security must constrain capability as well as behavior.
An agent that can effectively approve its own authority is not an enterprise control system. It is privileged access with a language model attached.
OpenShell is important because it treats the agent as an untrusted workload without making the agent useless. The workload can still reason, generate code, call services and adapt. What changes is that the final permission decision occurs somewhere the agent does not control.
That is likely to become a defining pattern of enterprise agent architecture whether or not NVIDIA owns the standard.
The second-order effect is potentially large. Identity teams will need nonhuman-agent identity models. Security teams will need policy and telemetry that understand tool calls and MCP traffic. Platform teams will need sandbox and workload lifecycle services. Application architects will have to decide where human approval remains mandatory. Audit systems will need to reconstruct not only what an agent said, but what it attempted, what was allowed, what was denied and under whose delegated authority.
The winners in enterprise agents may therefore be determined as much by control-plane architecture as by model intelligence.
That is the insider signal in NVIDIA's announcement. The race is no longer only to build smarter agents. It is to own the layer that decides how much power those agents are allowed to have.
Three moves for CIOs
-
— Define an agent execution boundary. Require production agents with credentials, write access, code execution or unattended authority to run inside an independently enforced runtime boundary rather than relying only on prompts, model safeguards or agent-harness permissions.
- Decision trigger: Before any agent receives persistent credentials, production write access or unattended execution authority.
- Why now: OpenShell makes the external-enforcement pattern concrete enough to test against real enterprise agents, and waiting until high-agency systems are already deployed will turn containment into retrofit work.
-
— Extend zero trust to nonhuman agents. Give each autonomous workload a verifiable identity, explicitly delegated authority, narrowly scoped service access and credentials that remain outside the agent whenever possible. Treat permission expansion as a governed event, not an agent convenience.
- Decision trigger: When agents begin crossing trust boundaries into enterprise APIs, databases, SaaS applications, infrastructure or other agents.
- Why now: IBM's integrations show identity, secrets and hybrid-cloud controls converging with agent runtimes. The architectural seams are being established now, before most enterprises have standardized them.
-
— Run an agent-containment bake-off before choosing the control plane. Test OpenShell alongside competing isolation and governance approaches using representative internal agents. Measure escape resistance, API-level policy expressiveness, credential exposure, policy-change workflow, multi-agent composition, audit quality, operational overhead and portability across models and infrastructure.
- Decision trigger: Before standardizing an enterprise runtime for high-agency production workloads or accepting a vendor's agent stack as the default governance layer.
- Why now: The architectural principle is maturing faster than the market. Enterprises have a short window to establish portable requirements before today's implementation choices harden into tomorrow's agent-control-plane lock-in.
Sources
- NVIDIA — Open Agent Safety Platform launch 1
- NVIDIA — Add Runtime Controls to AI Agents with OpenShell 2
- NVIDIA — How to Govern Autonomous Agents in Enterprise AI Factories 2
- IBM — Building Trust Into the Next Generation of AI Agents 3
- Reuters — independent reporting on NVIDIA's agent-safety launch 4
[^nvidia-launch][^nvidia-openshell][^reuters]