Sorami’s Hidden Network study found that a pod in an unrelated namespace could reach vendor-default Ray control ports in one EKS cluster, speaking cleartext gRPC. To close them, apply an ingress default-deny NetworkPolicy first. Then check Ray ports from inside the pod, encrypt traffic between nodes, and upgrade vLLM before trusting its API key.
Raw commands and outcomes are in Sorami’s public data repo. This guide separates measured controls from untested refinements.
Key takeaways
- Ray GCS and raylet ports were reachable from an unrelated namespace, in cleartext.
- An ingress default-deny NetworkPolicy blocked every Ray port that was open.
- Default scanners did not inspect the RayCluster, so its Ray pods were not analysed.
- Cilium plus WireGuard carried streaming, with a 3% to 8% difference in one run.
- A vLLM API key returning 401 does not show that the API is protected.
What did the study find?
The study deployed KubeRay 1.7.1 and Ray 2.52.0 with vLLM using defaults. From an unprivileged pod in an unrelated namespace, Ray GCS on port 6379 and raylet RPC on ports 10002 to 10006 answered in cleartext.
Four scanner outputs did not name the RayCluster; this supports an inference that its pod templates were not analysed, not a general scanner failure. Of 17 Ray sockets, 15 were absent from containerPort metadata. Undeclared ports are not blocked. The Ray security documentation says Ray should only run inside a trusted network.
How do you close the Ray and vLLM network?
Check CNI enforcement first. Sorami recorded the VPC CNI policy agent enabled. A policy object alone is not enforcement.
1. Apply an ingress default-deny NetworkPolicy
Deny ingress first, then add the tested allows. Engine-labelled pods may reach each other on all ports. The harness allowed port 8000 from all pods in two namespaces, not individual named callers. Combine namespace and pod selectors if you need a narrower caller rule; that refinement was not tested.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-ingress
namespace: inference
spec:
podSelector: {}
policyTypes: ["Ingress"]
It blocked seven previously open Ray listeners. Applying it live forced a restart. Install it before starting the engine. This ingress-only test did not measure egress containment or establish overhead. The policies are in the test harness. On EKS, consider NETWORK_POLICY_ENFORCING_MODE=strict so new pods start default-deny. We did not test the startup window of newly created pods.
2. Lock down the Ray Job API
The Ray Job API on 8265 accepts executable entrypoints by design. It answered on the default CPU image, not the GPU pipeline. No successful job submission or RCE was demonstrated. Authenticate it or disable the dashboard. Evaluate Ray token authentication (RAY_AUTH_MODE=token), which is off by default. We did not test it.
3. Find ports from the running pod
List listening sockets inside each Ray pod, or scan from a pod in another namespace against your own targets.
4. Encrypt traffic between nodes
The tested Linkerd sidecar configuration returned HTTP 503 for every request. The root cause was not isolated; this is not a general Linkerd verdict. Cilium WireGuard chained behind the VPC CNI carried 512 streaming requests; WireGuard alone was not isolated. One bracketed run showed 3.4% to 6.5% lower throughput and 3.2% to 8.4% higher median full-request latency. The closing baseline still had Cilium installed, so this does not isolate encryption overhead. When removing Cilium, set cni.uninstall=true. The default left a stale CNI config that broke new pods.
5. Upgrade vLLM, then set an API key
Set --api-key from a Kubernetes Secret on vLLM 0.22.0 or later. Versions from 0.3.0 up to 0.22.0 are affected by CVE-2026-48746, a Host header bypass. The normal key path returned 401 on two /v1 routes. That is not full authentication. Test every route, including /metrics and /tokenize, with no key and a wrong key.
Sorami’s view
This section is opinion. These are priorities, not measured implementation times.
Treat the Ray network as plumbing that no other workload should reach. Prioritise the measured ingress boundary, then verify it from an unrelated namespace. An API key sent in cleartext on a flat network is a weak layer on its own. A reachability test from a neighbour pod is part of our AI production readiness review and our cloud penetration testing.
Checklist
- Apply an ingress default-deny NetworkPolicy to every Ray namespace.
- Allow only named callers to reach the vLLM API port.
- Disable or authenticate the Ray Job API on port 8265.
- List listening sockets inside each Ray pod.
- Encrypt inter-node pod traffic with WireGuard or IPsec.
- Upgrade vLLM to 0.22.0 or later before setting an API key.
- Test every vLLM route with no key and a wrong key.
Related reading
The full Hidden Network report has the method, tables and limits. Our AI on Kubernetes Helm chart research covers single-chart defaults. For agents, see how an agent used DNS to get around its sandbox. More is in the guides index.
Sources
- Sorami: The Hidden Network: Ray Control-Plane Exposure in Distributed LLM Inference on Kubernetes (29 September 2026)
- Sorami: Hidden Network data, logs and test harness (GitHub)
- Ray Project: Security (Ray documentation)
- Ray Project: Ray token authentication
- GitHub Advisory Database: GHSA-94f4-hr76-p5j6 / CVE-2026-48746 (16 June 2026)
- Kubernetes documentation: Network Policies
- Amazon EKS User Guide: Configure network policy
- Cilium documentation: WireGuard Transparent Encryption
- Cilium documentation: Helm Reference (cni.uninstall)