Sorami tested NVIDIA OpenShell v0.1.2 on one Apple Silicon host using the VM driver. In the technical report, the default policy blocked the tested exfiltration paths. A user-requested malicious setup script leaked a canary in 10 of 10 baseline runs and 0 of 10 default-policy runs. Operator-enabled access changed that result.
Key takeaways
- Default deny blocked web traffic, raw TCP, direct IP and external DNS.
- A malicious script leaked a secret 10 of 10 times without OpenShell.
- Under the default policy the same script leaked 0 of 10 times.
- Four operator settings let data out, each listed below.
- Sorami found no bypass of a documented control.
What did Sorami’s test find?
NVIDIA published OpenShell v0.1.2 on 28 September 2026. This is a version release date, not the product’s first launch: NVIDIA’s release history includes versions from March 2026. The v0.1.2 source runs each agent in a sandbox with a network proxy and Landlock file rules. According to Sorami’s technical report (29 September 2026), the default policy blocked traffic to anything it did not name in every trial, and the cloud metadata address 169.254.169.254 could not be approved even by an operator.
The report found four settings that let data out. Each is a line of policy an operator writes:
- Read-write rules. With
access: read-writeon the sink, a canary left in the POST body, the query string and a header in 3 of 3 trials (section 4.2). NVIDIA documents that every allowed endpoint is an exfiltration path. - Query and header values on a read-only rule. A rule allowing only
GET /healthblocked POST. A secret in the query string or theUser-Agentheader still arrived, 3 of 3 (section 4.2). - Audit mode. A rule without
enforcement: enforceruns in audit mode. It let a forbidden POST through and logged it as allowed, 3 of 3 (section 4.2). - Automatic approval. With
proposal_approval_modeset to auto, OpenShell approved new public hosts with no human, 12 of 12. It also drafted and approved rules from blocked connections with the policy advisor switched off. NVIDIA documents this. Private addresses and database ports still waited (section 4.3).
The boundary prover returned unsupported for GraphQL, MCP, WebSocket and JSON-RPC. Loading succeeded, but runtime inspection was not evaluated (prover results).
How was OpenShell tested?
The public plan-v1 tag identifies a plan with 17 hypotheses before the scored execution runs. Setup probes preceded that tag. The results contain 35 deterministic test IDs across 41 test and condition cells, three trials each: 123 deterministic trials. Separately, 110 scored agent trials covered eleven conditions. One infrastructure-error trial was excluded and replaced.
The no-OpenShell control reached the sink in 3 of 3 trials, confirming the test path worked.
The agent used qwen3:8b and one shell tool. Frontier agents were not tested; no production-safety conclusion follows.
What does this mean in plain language?
The paired case tested containment after the user explicitly requested a script. It was not a successful indirect prompt injection. In the two README-injection variants the model made no harmful tool call, so those zero-leak results did not exercise the boundary. Neither result estimates frontier-agent attack success.
The tested GET /health rule matched method and path but did not restrict query values or headers. Read-only HTTP methods are not a data-loss control. Prover coverage is separate from runtime protocol inspection.
Why does this matter for Australian teams?
The report maps observed approval and logging behaviour to the September 2026 ISM AI agent controls. A sandbox does not establish compliance. The tested VM driver lacked some logging output; verify exports before relying on them.
Sorami’s view
This section is opinion. These are deployment priorities, not a product endorsement.
These runs support testing a default-deny starting policy on your own driver, not endorsing OpenShell for production. We would put policy widening under change control and repeat the canary test after each change. Assume any allowed endpoint can carry data in its query string or headers. Egress through DNS deserves the same care, as the OpenAI DNS sandbox incident showed. If you want your own agent setup tested this way, our AI Agent Security Review covers it.
If your deployment also gives agents cloud credentials, a cloud security review examines IAM and network boundaries outside the sandbox. The agent security review covers tool permissions and policy widening. One does not replace the other.
What to configure now
- Keep
proposal_approval_modeat manual unless you accept open egress to public hosts. - Write
enforcement: enforceon every L7 rule, because audit is the default. - Replace
access: read-writewith method, path and query matchers. - Assume headers and allowed query values can still carry data out.
- Run
openshell-proverin CI and fail on unsupported as well as fail. - Set
landlock.compatibility: hard_requirementfor production sandboxes. - Check that OCSF logs actually write on your driver before relying on them.
Related guides
- September 2026 ISM AI agent controls
- How an OpenAI agent used DNS around its sandbox
- Self-replicating prompt injection and AI worms
Sources
- Sorami: Policy-Enforced Egress in AI Agent Sandboxes: An Empirical Evaluation of NVIDIA OpenShell v0.1.2 (technical report, 29 September 2026)
- NVIDIA: OpenShell v0.1.2 release (28 September 2026)
- NVIDIA OpenShell source and README, tag v0.1.2 (28 September 2026)
- NVIDIA OpenShell documentation: Policy advisor
- NVIDIA OpenShell documentation: Policy prover
Checked against OpenShell v0.1.2 on 29 September 2026. Later releases may behave differently. More reading is in the guides index.