Skip to main content

We tested NVIDIA's new AI agent sandbox (OpenShell). The defaults held. Four settings let data out.

0
Medium
Published: 09/29/2026 (09/29/2026, 09:34:13 UTC)
Source: Reddit Cybersecurity

Description

NVIDIA OpenShell is an open-source sandbox for AI agents designed to block unauthorized network, file, and process access. In tests of version 0.1.2, the default policy effectively prevented data exfiltration, blocking all attempted leaks. However, four specific operator-configurable settings can allow data to leave the sandbox: read-write access rules, GET-only rules that still transmit data in query strings and headers, audit mode rules that log but do not block, and automatic approval of new public hosts without human intervention. The sandbox does not validate certain rule types (GraphQL, MCP, WebSocket, JSON-RPC), accepting them without warning. The sandbox is effective under default settings but requires careful policy management to maintain security.

Reddit Discussion

r/cybersecurity·posted by u/No-Peanut-6988
00

OpenShell is NVIDIA's open-source sandbox for AI agents: it blocks network, file and process access unless a policy allows it. We ran 123 trials on v0.1.2.

The good: every documented control held, and a malicious script run by an agent leaked a secret 0/10 times under the default policy (10/10 without OpenShell).

The four settings to check before you trust it:

  1. Read-write rules let data out.
  2. GET-only rules still carry data in query strings and headers.
  3. Rules in audit mode log but do not block.
  4. Auto-approval granted new public hosts with no human in 12/12 trials.

Plain-language write-up and checklist: https://sorami.com.au/guides/we-tested-nvidia-openshell/

Full report and logs: https://sorami.com.au/research/nvidia-openshell-agent-sandbox-test/

Affected software

Affected versions
=0.1.2

AI-Powered Analysis

Machine-generated threat intelligence

AILast updated: 09/29/2026, 09:47:51 UTC

Technical Analysis

NVIDIA OpenShell v0.1.2 runs AI agents in a sandbox that enforces network and file access policies using a default deny approach. Sorami's empirical evaluation showed that the default policy blocked all tested exfiltration attempts, with a malicious script leaking secrets 10/10 times without OpenShell and 0/10 times with it. Four operator-controlled policy settings can permit data exfiltration: read-write access rules allow data in POST bodies, query strings, and headers; GET-only rules still transmit data in query strings and headers; audit mode rules log but do not enforce blocking; and automatic approval mode grants new public hosts without human review. The policy prover does not support GraphQL, MCP, WebSocket, and JSON-RPC rules, yet the loader accepts them silently. The sandbox requires careful configuration, especially keeping proposal approval manual and enforcing all Layer 7 rules, to prevent data leaks.

Potential Impact

Under the default policy, OpenShell effectively prevents data exfiltration by AI agents, blocking all tested attempts to leak secrets. However, enabling certain policy settings can allow data to be exfiltrated, potentially exposing sensitive information. Automatic approval of new public hosts without human oversight can increase risk by allowing unvetted network access. The acceptance of unsupported rule types without warning could lead to misconfigurations and unintended access. Overall, the sandbox provides a strong boundary but relies heavily on correct operator policy configuration to maintain security.

Defensive Guidance

Use the default deny policy and keep proposal_approval_mode set to manual to prevent automatic approval of new hosts. Enforce all Layer 7 rules rather than using audit mode to ensure blocking rather than just logging. Avoid using read-write access rules unless absolutely necessary, and be aware that GET-only rules can still transmit data in query strings and headers. Run the policy prover in continuous integration to detect unsupported rules and fail builds accordingly. Treat any allowed endpoint as a potential exfiltration path and apply change control to all policy modifications. Verify that logging mechanisms are functioning correctly before relying on them.

Pro Console: star threats, build custom feeds, automate alerts via Slack, email & webhooks.Upgrade to Pro

Technical Details

Source Type
reddit
Subreddit
cybersecurity
Reddit Score
0
Discussion Level
minimal
Content Source
reddit_link_post
Post Type
link
Newsworthiness Assessment
{"score":27,"reasons":["external_link","established_author","very_recent"],"isNewsworthy":true}
Has External Source
true
Trusted Domain
false

Threat ID: 6abb893df7a7c54106275069

Added to database: 09/29/2026, 09:47:41 UTC

Last enriched: 09/29/2026, 09:47:51 UTC

Last updated: 09/29/2026, 18:36:24 UTC

Views: 14

Community Reviews

0 reviews

Crowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.

Sort by
Loading community insights…

Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.

Need more coverage?

Upgrade to Pro Console for AI refresh and higher limits.

For incident response and remediation, OffSeq services can help resolve threats faster.

Latest Threats

Breach by OffSeqOFFSEQFRIENDS — 25% OFF

Check if your credentials are on the dark web

Instant breach scanning across billions of leaked records. Free tier available.

Scan now
OffSeq TrainingCredly Certified

Lead Pen Test Professional

Technical5-day eLearningPECB Accredited
View courses