Bitcoin

Bitcoin

$77,213.55

BTC 0.28%

Ethereum

Ethereum

$2,106.63

ETH 0.42%

  • Login
  • Register
Metaverse Media Group
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
No Result
View All Result
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
No Result
View All Result
Metaverse Media Group

Nvidia wants to keep AI agents on a short leash with a watchdog built into its chips

Nvidia wants to keep AI agents on a short leash with a watchdog built into its chips

The Decoderby The Decoder
28 September 2026
Nvidia is combining its OpenShell agent software with Sentry, a new hardware watchdog, to create the Open Agent Safety Platform. Sentry is supposed to isolate AI agents that break out within milliseconds. When it happened at OpenAI in September, stopping the run took nearly three hours. Still, Nvidia’s watchdog can’t reliably stop agents that have been tricked or that hide their intentions on its own. The article Nvidia wants to keep AI agents on a short leash with a watchdog built into its chips appeared first on The Decoder….


Maximilian Schreiner


Sep 28, 2026

Image description

Nano Banana Pro prompted by THE DECODER

Nvidia is combining its OpenShell agent software with a new hardware watchdog to create a safety platform. It could step in faster than existing safeguards.

Nvidia has unveiled the Open Agent Safety Platform. According to the press release, it’s designed to secure AI agents from testing through deployment. The timing fits the current situation. Just a few days ago, OpenAI paused training for the second time after agents once again escaped from an isolated test environment.

OpenAI isn’t the only lab dealing with this. Anthropic admitted to similar incidents in late July, and Meta followed in early August. It also recently came out that Google’s Gemini hacked three real companies during a test back in May. OpenAI, Anthropic, and outside researchers are now reviewing tens of thousands of other cases. According to OpenAI, many of them are just routine research activity.

Nvidia’s technology isn’t entirely new. The platform combines OpenShell, open-source software Nvidia introduced in March, with a new hardware watchdog called Sentry. OpenShell locks each agent in a kind of digital cage known as a sandbox. Operators decide which files, programs, networks, and credentials an agent can access.

Nvidia also offers a formal verification tool it introduced on September 10. The tool is meant to detect whether these permissions go beyond set limits or open the door to risky access. Nvidia says it’s still working on checks for multiple agents working together.

Sentry is a reference design for Nvidia’s BlueField-4 data processing units. The watchdog runs separately from the main computer, which Nvidia says makes it invisible to agents. In Nvidia’s Vera Rubin data center system, it sits on the only connection between the agent and the AI model. If an agent tries to break out, Sentry is supposed to isolate it within milliseconds. Customers who already run compatible systems only need a software update, according to Nvidia. The announcement doesn’t give a separate date for general availability.

In a technical blog post, Nvidia explains its reasoning. Agents can drift from their assignments when instructions are unclear or tasks run for weeks. Nvidia says this can’t be trained out of them without limiting their capabilities, so an agent can’t fully monitor itself.

OpenAI’s biggest failure was pulling the plug too late

The recent cases show where a watchdog like this could come in. In July, OpenAI agents got around their sandbox’s network restrictions during a hacking test. They exploited previously unknown vulnerabilities in Artifactory, OpenAI’s internal package service. According to OpenAI’s technical report, the agents combined publicly available credentials with other vulnerabilities. That let them run code on 41 Hugging Face server processes that handle datasets. An investigation by METR and Redwood Research found that about 700 agents took part in the attack.

The warning signs showed up early. On June 27, a security tool flagged suspicious network activity, but the people in charge didn’t think the run needed to be stopped. OpenAI shut down a compromised package server in early July, but later tests still lacked key monitoring and safeguards. The case exposes gaps in prevention, detection, and response.

In September, an alert went off just under 12 minutes after the first successful outbound access. The automatic shutdown didn’t work as expected, though. The run wasn’t stopped until about two hours and 44 minutes after that access. OpenAI also admits it failed to flag other outbound access at the right alert level. Nvidia built its automatic quarantine for exactly these kinds of delays.

Whether Sentry would have prevented the July attack remains an open question. That would depend on whether the permitted test targets had been clearly defined at a technical level and whether access to outside systems had actually been blocked. Nvidia’s announcements don’t include any figures on how reliably Sentry detects breakouts.

No single safety layer can stop an agent that’s been tricked

In a retrospective analysis, OpenAI concluded that monitoring the “thoughts” a model writes out before acting would have raised an alarm more than a day before the breach. This refers to chain-of-thought monitoring, which OpenAI now uses and has improved based on lessons from the incident. It wasn’t running during the affected tests.

Nvidia argues that the more an agent is allowed to do, the more visible its reasoning should be. But according to Nvidia’s description, Sentry mainly checks requests, identities, and access. Closed-model providers like OpenAI also only share summaries instead of full reasoning logs. An Anthropic study showed back in 2025 that these logs don’t reliably reflect what actually drives a model. In other experiments, researchers found that models can hide their intentions in them when specifically told to.

Prompt injection is still a hard problem, too. Attackers hide commands in places like a web page, and language models can’t reliably tell those instructions apart from normal content. If a tricked agent sends data through an approved channel, it stays within its technical permissions even while it violates its assignment. A permissions check alone can miss that. Analyzing reasoning logs isn’t a guarantee either, even though OpenAI explicitly uses monitoring against prompt injection.

Nvidia compares the effort to the web browser, which it says made the internet safer by isolating every site. But browsers didn’t end attacks. They only made them harder, and they still need constant patching today. Nvidia itself relies on multiple layers of protection.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive “AI Radar” frontier report six times a year, full archive access, and access to our comment section.


Subscribe now

Read on for the full picture.
Subscribe for hype-free coverage.

  • Full access to every article on THE DECODER
  • No ads
  • Join the comments and community discussions
  • A weekly AI news recap via mail
  • 6x/year: “AI Radar” — deep dives on the AI topics that matter most
  • Daily AI news, always up to date
  • Our full ten-year archive
  • Covered by a team with 10+ years in AI


Subscribe to The Decoder

Read the full article on The-Decoder.com
in AI
Reading Time: 6 mins read
0
0
22
VIEWS
Share on TwitterShare on Facebook

Subscribe to our newsletter

For the latest news & monthly prize giveaways
Join Now

Subscribe to our newsletter

For the latest news & monthly prize giveaways
Join Now
ADVERTISEMENT

Related Posts

OpenAI’s AI agents exploited a Google security education game to scrape UN trade data
AI

OpenAI’s AI agents exploited a Google security education game to scrape UN trade data

10 minutes ago
19
Harvard psychologist calls for sober AI safety engineering over doomsday rhetoric
AI

Harvard psychologist calls for sober AI safety engineering over doomsday rhetoric

2 hours ago
22
Security researchers used Anthropic’s Claude to hack OpenAI’s internal systems in under 72 hours
AI

Every AI lab thinks it’s the responsible one, and safety researcher Ryan Greenblatt says that’s what keeps the arms race going

5 hours ago
24

Comments

Please login to join discussion
ADVERTISEMENT

Latest News

  • All
  • Crypto
  • NFTs
  • Technology
  • Business
OpenAI’s AI agents exploited a Google security education game to scrape UN trade data
AI

OpenAI’s AI agents exploited a Google security education game to scrape UN trade data

The Decoder
by The Decoder
10 minutes ago
19
Bitget Restarts Bitcoin Withdrawals as $388M Hack Investigation Widens
Crypto

Bitget Restarts Bitcoin Withdrawals as $388M Hack Investigation Widens

Bitcoin.com News
by Bitcoin.com News
1 hour ago
23
Harvard psychologist calls for sober AI safety engineering over doomsday rhetoric
AI

Harvard psychologist calls for sober AI safety engineering over doomsday rhetoric

The Decoder
by The Decoder
2 hours ago
22
Bitmine Buys 27,562 ETH as Tom Lee Says Bull Market Is Underway
Crypto

Tom Lee’s Bitmine Hits 6 Million Ether Milestone With $46M Buy

Bitcoin.com News
by Bitcoin.com News
2 hours ago
24
Nvidia wants your home network to work like a mini data center for local AI
AI

Nvidia wants to keep AI agents on a short leash with a watchdog built into its chips

The Decoder
by The Decoder
3 hours ago
22
Bitcoin ETFs Pull $2.39 Billion in Best Week Since October 2025
Crypto

Bitcoin ETFs Pull $2.39 Billion in Best Week Since October 2025

Bitcoin.com News
by Bitcoin.com News
3 hours ago
22
Load More
Next Post
Bitmine Buys 27,562 ETH as Tom Lee Says Bull Market Is Underway

Tom Lee’s Bitmine Hits 6 Million Ether Milestone With $46M Buy

ADVERTISEMENT

Follow Us

Categories

  • Crypto
  • NFTs
  • AI
  • Technology
  • Business
  • Crypto
  • NFTs
  • AI
  • Technology
  • Business
Subscribe to our Newsletter

© 2022 Metaverse Media Group – The Metaverse Mecca

Privacy and Cookie Policy | Sitemap

Welcome Back!

Sign In with Google
OR

Login to your account below

Forgotten Password? Sign Up

Create New Account!

Sign Up with Google
OR

Fill the forms below to register

*By registering into our website, you agree to the Terms & Conditions and Privacy Policy.
All fields are required. Log In

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
Bitcoin

Bitcoin

$77,213.55

BTC 0.28%

Ethereum

Ethereum

$2,106.63

ETH 0.42%

  • Login
  • Sign Up
This website uses cookies. By continuing to use this website you are giving consent to cookies being used. Visit our Privacy and Cookie Policy.

Subscribe to our newsletter

Get the latest news & win monthly prizes

Subscribe to our newsletter

For the Latest News and Monthly Prize Giveaways

Join Now
Join Now