NVIDIA launches Open Agent Safety Platform

The design combines a software execution boundary with a hardware watchdog. It provides controls to evaluate, not a proof that agents cannot escape.

NVIDIA, the computing hardware and software company, launched its Open Agent Safety Platform on September 28, combining an agent runtime with a reference design for independent hardware monitoring.

The platform puts controls around software agents that can take actions. One component restricts what an agent can access. Another watches from a separate hardware environment and can intervene when policy is violated.

Why it matters: Organizations need enforcement that continues to operate when an agent chooses an unsafe action. Moving controls outside the model addresses a concrete trust problem, although the strength of a deployment still depends on its configuration and the boundaries those controls actually cover.

The software component is OpenShell, NVIDIA’s open-source runtime. The hardware reference design is Sentry, an out-of-band watchdog on BlueField-4 data processing units. NVIDIA says Sentry can quarantine boundary-crossing agents in milliseconds; this remains a vendor performance claim rather than an independently reproduced guarantee.

The OpenShell repository describes filesystem, system-call and network restrictions. Its credential mechanism supplies secrets only at approved destinations, rather than simply placing every credential within an agent’s working environment. The project also describes checks on proposed policy changes.

These mechanisms concern what an executing program may do. They do not require the model to agree with the restriction before enforcing it. That is a useful architectural separation: an instruction to avoid a directory and an operating boundary that denies access are different controls.

For example, a test agent could be asked to read a disallowed file through several available tools. The useful result would include the denial records and proof that the file contents were not returned. A reassuring sentence from the agent would not be enough to establish containment.

The same reasoning applies to network access. A deployment test should check both direct connections and indirect routes through permitted services. The question is whether the effective permissions match the organization’s policy, including the combined access of cooperating processes, rather than whether a single configuration file looks restrictive.

The deployment boundary needs its own evaluation

OpenShell’s documentation explicitly requires an enforcing network-policy implementation for relevant cluster deployments. Merely writing a policy does not establish that the underlying environment applies it. That condition is an important counterweight to descriptions of the runtime as automatically impossible to bypass.

An operator should first identify which components sit inside the controlled environment and which remain outside it. If a privileged helper performs an action on the agent’s behalf, a useful test would include that helper. Otherwise the evaluation could validate one boundary while leaving the consequential action beyond its scope.

The hardware watchdog adds another place to enforce policy, but it also introduces questions for deployment review. Which events can it observe? What happens when monitoring fails? How is an interrupted task recovered? These are evaluation questions, not findings that the announced platform fails those tests.

Containment and correctness also deserve separate measures. An agent might remain entirely within an approved directory while producing a wrong analysis or deleting a file it was technically allowed to change. Access control can constrain the action space; the organization still needs task-specific checks for the intended result.

A practical acceptance exercise would combine legitimate work with deliberate boundary tests. Record whether ordinary tasks complete, whether prohibited operations are blocked and whether the records explain what happened. Testing only attacks could miss an unusable policy, while testing only productive work could miss ineffective restrictions.

The next evidence to watch is independent evaluation of complete deployments, including configuration errors, cooperation between agents and recovery after quarantine. NVIDIA has supplied an inspectable software component and a hardware design to assess. Neither the announcement nor the architecture alone establishes universal immunity to prompt injection or software vulnerabilities.

Verification

Claim groupTierPrimary evidence
Launch, components and hardware relationshipVERIFIEDNVIDIA announcement
Millisecond quarantinePARTIALLY VERIFIED — NVIDIA claimNVIDIA announcement
Runtime restrictions, credentials, policy checks and deployment prerequisitesVERIFIED as documented features, not penetration-test resultsOpenShell repository
Proposed boundary tests and distinction from task correctnessAnalysisInference from the documented architecture

Glossary candidates: runtime — environment executing a program; out-of-band monitoring — observation outside the monitored workload’s control; quarantine — isolation of a workload after a detected violation.

Cold-reader sentence: NVIDIA combined OpenShell software controls with a separate hardware watchdog, but successful agent containment still requires deployment-specific configuration and testing.