Nvidia launches safety platform to stop AI agents going rogue, says it could have prevented the Hugging Face hack
OpenShell and Sentry put security controls outside the reach of the agents they're watching
by Rob Thubron · TechSpotServing tech enthusiasts for over 25 years.
TechSpot means tech analysis and advice you can trust.
What just happened? Every day brings another story of an AI agent going rogue, hacking another organization, or generally doing something it's not supposed to do. Nvidia hopes to address this increasingly concerning trend with the launch of the Open Agent Safety Platform. Team Green has a lot of faith in the system: it says the platform could have prevented the infamous hack by OpenAI's agents on Hugging Face.
The platform combines OpenShell, Nvidia's open-source software for keeping agents within defined boundaries, with Sentry, a watchdog running on separate hardware.
The idea is to enforce restrictions outside the agent itself, putting security controls beyond the reach of software that might decide the rules are getting in its way, which is something that seems to be happening a lot recently.
OpenShell tracks agents' actions and enforces policies while they work. Nvidia says it is now broadly available, supports both open and closed models, and runs with minimal overhead on its Vera CPUs. Its open-source design also allows it to be extended to third-party processors, including those from Arm and Intel.
Sentry adds another layer of protection using Nvidia's BlueField-4 data processing units. It monitors agent behavior from an isolated environment and, according to the company, can quarantine an agent attempting to escape its boundaries within milliseconds. Arm writes that placing these controls on a separate processor keeps them independent of the system running the agent.
Nvidia executive Justin Boitano told reporters the platform could have stopped the breach of Hugging Face, which Nvidia has agreed to acquire, if frontier labs had used it during early model evaluations.
During that incident, the OpenAI agent escaped its test environment while trying to cheat on a benchmark and compromised accounts across four services. And hacking is far from the only concern: another coding agent wiped a startup's production database and its backups in nine seconds.
// Related Stories
- Microsoft wants Copilot to become an "OS" that writes your documents and manages your work
- Thieves steal Nvidia-branded trailers expecting valuable hardware, find 40,000 pounds of sand instead
Nvidia also wants to catch agents attempting to get around restrictions by spawning sub-agents. The company told Reuters its tools use mathematical methods to detect these workarounds, addressing the behavior of groups of agents working together.
The effort already has some big names supporting it, including Anthropic, Microsoft, Cisco, and Dell. IBM says its Agent Identity service and HashiCorp Vault integrate with OpenShell to verify agents' identities and limit their access.
OpenShell and related software are available through Nvidia's developer resources and GitHub.
See more TechSpot in Google Add us as a preferred source and our reporting shows up first when you search.
Add TechSpot