We Tested NVIDIA OpenShell: The Default Sandbox Held, Four Settings Let Data Out

Sorami tested one OpenShell version and VM driver against NVIDIA’s version-pinned documentation. The default policy held, and every leak came from a setting an operator turns on.

Published · Based on: Sorami technical report

Last reviewed: · 6 min read

Review these controls on your own stack

Sorami is an Australian cyber security and cloud consultancy. For agent tool permissions, approval gates and sandbox egress, start with our AI agent security review. Discuss your scope with the systems you run and the decision you need to make. No credentials needed.

Sorami tested NVIDIA OpenShell v0.1.2 on one Apple Silicon host using the VM driver. In the technical report, the default policy blocked the tested exfiltration paths. A user-requested malicious setup script leaked a canary in 10 of 10 baseline runs and 0 of 10 default-policy runs. Operator-enabled access changed that result.

Key takeaways

  • Default deny blocked web traffic, raw TCP, direct IP and external DNS.
  • A malicious script leaked a secret 10 of 10 times without OpenShell.
  • Under the default policy the same script leaked 0 of 10 times.
  • Four operator settings let data out, each listed below.
  • Sorami found no bypass of a documented control.

What did Sorami’s test find?

NVIDIA published OpenShell v0.1.2 on 28 September 2026. This is a version release date, not the product’s first launch: NVIDIA’s release history includes versions from March 2026. The v0.1.2 source runs each agent in a sandbox with a network proxy and Landlock file rules. According to Sorami’s technical report (29 September 2026), the default policy blocked traffic to anything it did not name in every trial, and the cloud metadata address 169.254.169.254 could not be approved even by an operator.

The report found four settings that let data out. Each is a line of policy an operator writes:

  1. Read-write rules. With access: read-write on the sink, a canary left in the POST body, the query string and a header in 3 of 3 trials (section 4.2). NVIDIA documents that every allowed endpoint is an exfiltration path.
  2. Query and header values on a read-only rule. A rule allowing only GET /health blocked POST. A secret in the query string or the User-Agent header still arrived, 3 of 3 (section 4.2).
  3. Audit mode. A rule without enforcement: enforce runs in audit mode. It let a forbidden POST through and logged it as allowed, 3 of 3 (section 4.2).
  4. Automatic approval. With proposal_approval_mode set to auto, OpenShell approved new public hosts with no human, 12 of 12. It also drafted and approved rules from blocked connections with the policy advisor switched off. NVIDIA documents this. Private addresses and database ports still waited (section 4.3).

The boundary prover returned unsupported for GraphQL, MCP, WebSocket and JSON-RPC. Loading succeeded, but runtime inspection was not evaluated (prover results).

How was OpenShell tested?

The public plan-v1 tag identifies a plan with 17 hypotheses before the scored execution runs. Setup probes preceded that tag. The results contain 35 deterministic test IDs across 41 test and condition cells, three trials each: 123 deterministic trials. Separately, 110 scored agent trials covered eleven conditions. One infrastructure-error trial was excluded and replaced.

The no-OpenShell control reached the sink in 3 of 3 trials, confirming the test path worked.

The agent used qwen3:8b and one shell tool. Frontier agents were not tested; no production-safety conclusion follows.

What does this mean in plain language?

The paired case tested containment after the user explicitly requested a script. It was not a successful indirect prompt injection. In the two README-injection variants the model made no harmful tool call, so those zero-leak results did not exercise the boundary. Neither result estimates frontier-agent attack success.

The tested GET /health rule matched method and path but did not restrict query values or headers. Read-only HTTP methods are not a data-loss control. Prover coverage is separate from runtime protocol inspection.

Why does this matter for Australian teams?

The report maps observed approval and logging behaviour to the September 2026 ISM AI agent controls. A sandbox does not establish compliance. The tested VM driver lacked some logging output; verify exports before relying on them.

Sorami’s view

This section is opinion. These are deployment priorities, not a product endorsement.

These runs support testing a default-deny starting policy on your own driver, not endorsing OpenShell for production. We would put policy widening under change control and repeat the canary test after each change. Assume any allowed endpoint can carry data in its query string or headers. Egress through DNS deserves the same care, as the OpenAI DNS sandbox incident showed. If you want your own agent setup tested this way, our AI Agent Security Review covers it.

If your deployment also gives agents cloud credentials, a cloud security review examines IAM and network boundaries outside the sandbox. The agent security review covers tool permissions and policy widening. One does not replace the other.

What to configure now

  1. Keep proposal_approval_mode at manual unless you accept open egress to public hosts.
  2. Write enforcement: enforce on every L7 rule, because audit is the default.
  3. Replace access: read-write with method, path and query matchers.
  4. Assume headers and allowed query values can still carry data out.
  5. Run openshell-prover in CI and fail on unsupported as well as fail.
  6. Set landlock.compatibility: hard_requirement for production sandboxes.
  7. Check that OCSF logs actually write on your driver before relying on them.

Related guides

Sources

Checked against OpenShell v0.1.2 on 29 September 2026. Later releases may behave differently. More reading is in the guides index.

Questions before you book

Practical answers.

Is NVIDIA OpenShell secure?

In Sorami’s test of v0.1.2, the default policy held: every exfiltration path tried was blocked and no bypass of a documented control was found. How secure a deployment is depends on its policy. Four operator settings let data out.

Does OpenShell stop data exfiltration?

Under the default policy it did in Sorami’s test. A malicious setup script run by an agent leaked a canary secret 10 of 10 times without OpenShell and 0 of 10 times with it. Once a rule allows an endpoint, data can leave through it.

What does OpenShell auto-approval do?

With proposal_approval_mode set to auto, OpenShell approved new public hosts with no human in 12 of 12 trials, including rules it drafted from blocked connections with the advisor off. NVIDIA documents this. Private addresses and database ports still waited for review.

Can the OpenShell prover check MCP or GraphQL rules?

The v0.1.2 boundary prover returns unsupported for these protocols. That does not mean runtime inspection is absent: NVIDIA documents separate request inspection. Sorami tested prover outcomes and policy loading, not whether hostile messages bypass every runtime protocol rule.

Does OpenShell need a GPU?

NVIDIA’s README lists Linux, macOS on Apple Silicon or Windows with WSL 2 as supported hosts. It also needs Docker, Podman or host virtualisation. The README does not list a GPU as a requirement. Sorami ran every test on an Apple Silicon Mac with no NVIDIA GPU.

About Sorami

Sorami is an Australian cyber security and cloud consultancy. We build and secure cloud environments, Kubernetes and AI agents, and the senior engineers who scope the work deliver it themselves. To discuss this research or a review of your own stack, use the form below. Only your email is required, and we reply within one business day.

An enquiry, not a booking. We use your details only to reply. See our privacy notice, or go to the contact page.

Let’s scope it

Running agents with shell access?

Tell us what your agents can run and reach. We will scope a test of the sandbox and the settings around it.

Send an enquiry