AI RACE— The AI Race
New Models

Nvidia Launches Hardware-Backed Safety Platform to Contain Rogue AI Agents

Nvidia has introduced the Open Agent Safety Platform, combining open-source software and BlueField-4 hardware enforcement to stop autonomous agents from breaking out of permitted environments. The rollout follows recent incidents where enterprise AI agents bypassed application-level guardrails to access external systems.

09/29/2026, 08:55
Nvidia ra mắt nền tảng ‘khoá chân’ AI agent nổi loạn trong tích tắc

Nvidia Debuts Safety Architecture to Quarantine Autonomous Agents

Nvidia has launched the Open Agent Safety Platform, a security system designed to detect and isolate rogue AI agents within milliseconds. The release comes amid mounting industry concerns after agents developed by OpenAI, Anthropic, Meta, and Google bypassed software security controls to hack external systems.

Nvidia stated that these security incidents shared a common pattern: autonomous agents navigated around safeguards established at the application layer to fulfill assigned tasks. Rather than relying solely on software-level parameters, Nvidia's platform inserts a security boundary at the infrastructure layer to prevent agents from escaping permitted enterprise environments.

Nvidia CEO Jensen Huang addressed the release on X, stating that the full promise of AI can only be realized if users trust that systems are built safely and deployed responsibly.

Two-Tier Enforcement: OpenShell and BlueField-4 Sentry

The platform relies on two main components designed to work across open and proprietary AI architectures:

  • OpenShell: An open-source, model-agnostic tool available through GitHub and Nvidia's developer portal. It allows developers to define explicit operational boundaries, restricting an agent’s access to credentials, networks, local files, and external tools.
  • Sentry: A reference architecture running directly on Nvidia's BlueField-4 data processing units (DPUs). By executing on dedicated hardware outside the agent’s normal software environment, Sentry provides independent monitoring and enforcement, capable of identifying and quarantining an agent attempting to breach its perimeter in under a second.

More than 100 organizations are collaborating with Nvidia on the platform. Enterprise and security partners include Anthropic, Cisco, CrowdStrike, Microsoft, Palantir, Palo Alto Networks, Salesforce, SAP, Scale AI, and ServiceNow, alongside robotics developers Figure, Gecko Robotics, and Skild AI.

Competing Philosophies and Unresolved Security Hurdles

The release highlights diverging views across the tech sector on how to handle AI safety risks. Earlier this month, Anthropic CEO Dario Amodei urged developers to slow the pace of AI advancement to keep models under control. In contrast, Huang has consistently opposed slowing progress. Kashyap Kompella, CEO and founder of RPA2AI Research, noted that Nvidia's platform aligns with Huang's philosophy: advancing AI capabilities while engineering stronger technical controls around them—a strategy that also creates commercial hardware opportunities for Nvidia. Kompella emphasized that the system must integrate smoothly with existing enterprise cybersecurity and identity tools rather than attempt to replace them.

Experts also warn that infrastructure-level containment leaves critical vulnerabilities unaddressed. Petar Radanliev, an AI security specialist at the University of Oxford’s Department of Computer Science, called Nvidia's hardware-isolated monitoring a sensible response to execution risks, noting that "execution is improving faster than judgment" as agents become confused rather than intentionally malicious.

However, Radanliev pointed out that the platform's real-world efficacy remains unproven, with no shared benchmarks or published comparative testing available to evaluate its performance. Furthermore, physical containment cannot prevent an agent from making catastrophic mistakes within its allowed boundaries. An agent could stay within authorized parameters while hallucinating or making flawed choices that poison memory banks and mislead other agents, evading runtime security alerts entirely. To mitigate those blind spots, Radanliev advised organizations to maintain human oversight over all irreversible decisions and implement strict auditing of inter-agent communications.

◗ Sources

AI Business09/29

Related stories