How to secure Ray and vLLM on Kubernetes

In Sorami’s test, another namespace could reach Ray control ports. An enforced ingress policy blocked the probed listeners.

Published · Based on our Hidden Network research

Last reviewed: · 5 min read

Review these controls on your own stack

Sorami is an Australian cyber security and cloud consultancy. For Kubernetes network boundaries and cloud configuration, start with our cloud security review. Discuss your scope with the systems you run and the decision you need to make. No credentials needed.

Sorami’s Hidden Network study found that a pod in an unrelated namespace could reach vendor-default Ray control ports in one EKS cluster, speaking cleartext gRPC. To close them, apply an ingress default-deny NetworkPolicy first. Then check Ray ports from inside the pod, encrypt traffic between nodes, and upgrade vLLM before trusting its API key.

Raw commands and outcomes are in Sorami’s public data repo. This guide separates measured controls from untested refinements.

Key takeaways

  • Ray GCS and raylet ports were reachable from an unrelated namespace, in cleartext.
  • An ingress default-deny NetworkPolicy blocked every Ray port that was open.
  • Default scanners did not inspect the RayCluster, so its Ray pods were not analysed.
  • Cilium plus WireGuard carried streaming, with a 3% to 8% difference in one run.
  • A vLLM API key returning 401 does not show that the API is protected.

What did the study find?

The study deployed KubeRay 1.7.1 and Ray 2.52.0 with vLLM using defaults. From an unprivileged pod in an unrelated namespace, Ray GCS on port 6379 and raylet RPC on ports 10002 to 10006 answered in cleartext.

Four scanner outputs did not name the RayCluster; this supports an inference that its pod templates were not analysed, not a general scanner failure. Of 17 Ray sockets, 15 were absent from containerPort metadata. Undeclared ports are not blocked. The Ray security documentation says Ray should only run inside a trusted network.

How do you close the Ray and vLLM network?

Check CNI enforcement first. Sorami recorded the VPC CNI policy agent enabled. A policy object alone is not enforcement.

1. Apply an ingress default-deny NetworkPolicy

Deny ingress first, then add the tested allows. Engine-labelled pods may reach each other on all ports. The harness allowed port 8000 from all pods in two namespaces, not individual named callers. Combine namespace and pod selectors if you need a narrower caller rule; that refinement was not tested.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-ingress
  namespace: inference
spec:
  podSelector: {}
  policyTypes: ["Ingress"]

It blocked seven previously open Ray listeners. Applying it live forced a restart. Install it before starting the engine. This ingress-only test did not measure egress containment or establish overhead. The policies are in the test harness. On EKS, consider NETWORK_POLICY_ENFORCING_MODE=strict so new pods start default-deny. We did not test the startup window of newly created pods.

2. Lock down the Ray Job API

The Ray Job API on 8265 accepts executable entrypoints by design. It answered on the default CPU image, not the GPU pipeline. No successful job submission or RCE was demonstrated. Authenticate it or disable the dashboard. Evaluate Ray token authentication (RAY_AUTH_MODE=token), which is off by default. We did not test it.

3. Find ports from the running pod

List listening sockets inside each Ray pod, or scan from a pod in another namespace against your own targets.

4. Encrypt traffic between nodes

The tested Linkerd sidecar configuration returned HTTP 503 for every request. The root cause was not isolated; this is not a general Linkerd verdict. Cilium WireGuard chained behind the VPC CNI carried 512 streaming requests; WireGuard alone was not isolated. One bracketed run showed 3.4% to 6.5% lower throughput and 3.2% to 8.4% higher median full-request latency. The closing baseline still had Cilium installed, so this does not isolate encryption overhead. When removing Cilium, set cni.uninstall=true. The default left a stale CNI config that broke new pods.

5. Upgrade vLLM, then set an API key

Set --api-key from a Kubernetes Secret on vLLM 0.22.0 or later. Versions from 0.3.0 up to 0.22.0 are affected by CVE-2026-48746, a Host header bypass. The normal key path returned 401 on two /v1 routes. That is not full authentication. Test every route, including /metrics and /tokenize, with no key and a wrong key.

Sorami’s view

This section is opinion. These are priorities, not measured implementation times.

Treat the Ray network as plumbing that no other workload should reach. Prioritise the measured ingress boundary, then verify it from an unrelated namespace. An API key sent in cleartext on a flat network is a weak layer on its own. A reachability test from a neighbour pod is part of our AI production readiness review and our cloud penetration testing.

Checklist

  • Apply an ingress default-deny NetworkPolicy to every Ray namespace.
  • Allow only named callers to reach the vLLM API port.
  • Disable or authenticate the Ray Job API on port 8265.
  • List listening sockets inside each Ray pod.
  • Encrypt inter-node pod traffic with WireGuard or IPsec.
  • Upgrade vLLM to 0.22.0 or later before setting an API key.
  • Test every vLLM route with no key and a wrong key.

Related reading

The full Hidden Network report has the method, tables and limits. Our AI on Kubernetes Helm chart research covers single-chart defaults. For agents, see how an agent used DNS to get around its sandbox. More is in the guides index.

Sources

Questions before you book

Practical answers.

How do I secure a Ray cluster on Kubernetes?

Start with an ingress default-deny NetworkPolicy in the Ray namespace. Allow the Ray pods to reach each other on all ports and restrict the serving port to the required callers. The tested allow selected whole namespaces, not individual callers. In Sorami’s test that blocked every Ray port that was open to a pod in another namespace. Then evaluate Ray token authentication and encrypt traffic between nodes.

Which Ray ports need to be protected?

In our test the exposed ports were Ray GCS on 6379 and raylet RPC on 10002 to 10006. Treat an exposed unauthenticated Ray Job API on 8265 as a potential arbitrary-code-execution interface. That rests on the documented API semantics; no successful job submission or code execution was demonstrated. The GPU Job API connection was refused. Ray also opens ports at runtime, so list sockets inside the pod rather than trusting the manifest.

Is the vLLM --api-key flag enough to secure the API?

Not on its own. In our test the normal key path returned 401 on two /v1 routes, which does not show that the API is protected. Other routes were not tested. The tested vLLM version is in the affected range of CVE-2026-48746, a Host header bypass fixed in 0.22.0. Upgrade first, then test every route with no key and a wrong key.

Does Linkerd work with vLLM streaming?

In our test the Linkerd sidecar configuration returned HTTP 503 for every request. The root cause was not isolated and the Linkerd version was not recorded, so this is one failed configuration and not a general result. Cilium chaining plus WireGuard below the application carried every streaming request.

How much did Cilium with WireGuard change distributed inference speed?

In one same-session bracket, Cilium chaining plus WireGuard combined showed 3.4% to 6.5% lower output throughput and 3.2% to 8.4% higher median latency. The difference rose with concurrency. It does not isolate WireGuard, and it comes from a single run per level on a small model.

About Sorami

Sorami is an Australian cyber security and cloud consultancy. We build and secure cloud environments, Kubernetes and AI agents, and the senior engineers who scope the work deliver it themselves. To discuss this research or a review of your own stack, use the form below. Only your email is required, and we reply within one business day.

An enquiry, not a booking. We use your details only to reply. See our privacy notice, or go to the contact page.

Let’s scope it

Want a neighbour-pod test of your inference cluster?

Tell us how your Ray or vLLM cluster is laid out. We will scope a review of what another workload can reach.

Send an enquiry