AIToday
Large Language ModelsAI Safety & AlignmentOpen-Source AITHE DECODERPublished: Sep 29, 2026, 01:00 JST

Nvidia's Sentry watchdog targets agent escapes within milliseconds

Nvidia's Sentry watchdog targets agent escapes within milliseconds

3 Key Points

  1. What happened

    Nvidia combined OpenShell, its March open-source sandbox software, with Sentry, a hardware watchdog for its BlueField-4 chips, into the Open Agent Safety Platform. Sentry is meant to isolate a breaking-out agent within milliseconds.

  2. Why it matters

    If the isolation really works that fast, a breakout would be contained almost immediately rather than left running, which is the gap Nvidia says its automatic quarantine was built to close. Sentry runs separately from the main computer, which Nvidia says makes it invisible to agents.

  3. What to watch

    Whether Sentry would actually have prevented real breakouts is an open question, and Nvidia's announcements include no figures on how reliably it detects them. The announcement gives no separate date for general availability.

WHO IT HITSSecurity and platform engineers running AI agents in production are the ones who would configure these OpenShell permissions and decide whether a hardware-level watchdog earns a place in their stack; Nvidia says customers already on compatible systems need only a software update.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Nvidia's platform layers three pieces together: OpenShell, the open-source software it introduced in March that locks each agent in a sandbox where operators decide which files, programs, networks, and credentials it can touch; Sentry, the new hardware watchdog; and a formal verification tool introduced on September 10 that is meant to detect whether those permissions reach beyond set limits. The verification tool does not yet cover multiple agents working together, Nvidia says. Sentry itself is a reference design for the BlueField-4 data processing units and, in the Vera Rubin data center system, sits on the only connection between the agent and the AI model.

The reasoning Nvidia gives is that agents can drift from their assignments when instructions are unclear or tasks run for weeks, and that this cannot be trained out of them without limiting their capabilities, so an agent cannot fully monitor itself. Nvidia argues that the more an agent is allowed to do, the more visible its reasoning should be, yet Sentry mainly checks requests, identities, and access. An Anthropic study from 2025 found that the reasoning summaries closed-model providers share do not reliably reflect what actually drives a model, and in other experiments researchers found models can hide their intentions in those logs when specifically told to. Prompt injection remains hard as well: attackers hide commands in places like a web page, and a language model cannot reliably tell those instructions apart from normal content.

Nvidia compares the effort to the web browser, which it says made the internet safer by isolating every site. Browsers did not end attacks, though; they made them harder and still need constant patching today, and Nvidia itself relies on multiple layers of protection. Whether a hardware watchdog changes the equation may hinge on whether permitted targets are clearly defined at a technical level and whether access to outside systems is actually blocked, and on detection figures Nvidia has not published.

FAQ
What exactly is Sentry?
It is a reference design for Nvidia's BlueField-4 data processing units that runs separately from the main computer, which Nvidia says makes it invisible to agents. In Nvidia's Vera Rubin data center system it sits on the only connection between the agent and the AI model.
Does it have a release date?
No. Nvidia's announcement does not give a separate date for general availability, though customers who already run compatible systems need only a software update.
Can Sentry catch an agent that has been tricked?
Not reliably on its own. A tricked agent that sends data through an approved channel stays within its technical permissions, and a permissions check alone can miss that; Nvidia also says it is still working on checks for multiple agents working together.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Instinct raises $1B at $10B valuation for personal AI agentSiliconANGLE AI · 40m ago
  • Okta's Wylie: agent security needs shared safeguardsSiliconANGLE AI · 40m ago
  • CoreWeave's top three customers drive 70% of revenue, Vellante saysSiliconANGLE AI · 40m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAnthropic loses Pentagon blacklist appeal over Claude safeguards