Skip to main content

The Hidden Network: Ray Control-Plane Exposure in Distributed LLM Inference on Kubernetes (Empirical Study, EKS)

0
Medium
Published: 09/29/2026 (09/29/2026, 09:50:18 UTC)
Source: Reddit NetSec

Description

An empirical study evaluated the network exposure of Ray control-plane components in distributed large language model (LLM) inference deployments on Kubernetes using Amazon EKS 1.35. The study found that multiple Ray control-plane sockets were not declared in pod metadata, allowing an unprivileged pod in a different namespace to access Ray GCS and raylet RPC services over unauthenticated cleartext gRPC. Default security scanners did not detect this exposure. A default-deny NetworkPolicy blocked access in steady state, and WireGuard encryption added modest overhead. The Ray Job API responded to unauthenticated requests on CPU Ray images but was not active in GPU deployments. The findings reflect vendor default configurations rather than new vulnerabilities, and a related CVE affecting vLLM was already fixed upstream.

Reddit Discussion

r/netsec·posted by u/No-Peanut-6988
00

We evaluated the runtime network surface of distributed LLM serving (KubeRay + vLLM pipeline parallelism) on an Amazon EKS 1.35 cluster from the perspective of an unprivileged neighbour pod in an unrelated namespace.

Summary of observations:

- 15 of 17 observed Ray listening sockets were omitted from pod declared containerPort metadata.

- An unprivileged neighbour pod in an unrelated namespace reached Ray GCS (6379) and raylet RPC (10002 to 10006) over unauthenticated cleartext gRPC.

- Four default security scanners (Trivy, Checkov, Kubescape, kube-linter) did not inspect the RayCluster custom resource, failing to represent runtime exposure.

- An ingress default-deny NetworkPolicy blocked all probed Ray ports from the neighbour pod in steady state.

- WireGuard encryption via Cilium chained with AWS VPC CNI carried the workload with an observed 3.4% to 6.5% throughput drop across a bracketed single-run test.

- The Ray Job API (8265) answered unauthenticated requests on default CPU Ray images, but was not listening in the GPU vLLM deployment (RCE against the GPU inference stack was not demonstrated).

All test manifests, scanner outputs, network logs, and repro scripts are open source:

https://github.com/Sorami-Consulting-AU/distributed-llm-inference-hidden-network

AI-Powered Analysis

Machine-generated threat intelligence

AILast updated: 09/29/2026, 10:32:54 UTC

Technical Analysis

This research analyzed the runtime network surface of distributed LLM serving using KubeRay and vLLM pipeline parallelism on an Amazon EKS 1.35 cluster. It discovered that 15 of 17 Ray listening sockets were omitted from pod containerPort metadata, enabling an unprivileged pod in an unrelated namespace to reach Ray GCS (port 6379) and raylet RPC ports (10002-10006) over unauthenticated cleartext gRPC. Four default Kubernetes security scanners failed to detect this exposure because they did not inspect the RayCluster custom resource. A default-deny NetworkPolicy effectively blocked these ports from neighbor pods in steady state. WireGuard encryption chained with AWS VPC CNI introduced a 3.4% to 6.5% throughput drop. The Ray Job API (port 8265) accepted unauthenticated requests on default CPU Ray images but was not listening in GPU vLLM deployments. The study's disclosures clarify that these behaviors are documented vendor defaults, not new vulnerabilities, and that a previously known vLLM API key bypass (CVE-2026-48746) has been fixed upstream.

Potential Impact

The study highlights that default vendor configurations for Ray control-plane components in Kubernetes clusters can expose unauthenticated gRPC services to unprivileged pods in unrelated namespaces. This exposure could allow unauthorized access to Ray control-plane services, potentially impacting confidentiality and integrity of distributed LLM inference operations. However, the presence of default-deny NetworkPolicies mitigates this exposure in steady state. The Ray Job API's acceptance of unauthenticated requests on CPU Ray images represents an additional risk vector, though no remote code execution was demonstrated in GPU deployments. The findings do not represent new vulnerabilities but rather documented default behaviors, with a related vLLM API key bypass already fixed upstream.

Defensive Guidance

The default-deny NetworkPolicy effectively blocks unauthorized access to Ray control-plane ports and should be maintained or implemented if absent. Vendors and operators should ensure that RayCluster custom resources are properly inspected by security scanners, as default tools may not detect this exposure. Enabling WireGuard encryption chained with AWS VPC CNI can provide additional network-layer protection with minimal throughput impact. Operators should update to patched versions of vLLM to address the previously known API key bypass (CVE-2026-48746). No new patches are indicated for the observed default behaviors, but hardening network policies and authentication configurations is recommended.

Pro Console: star threats, build custom feeds, automate alerts via Slack, email & webhooks.Upgrade to Pro

Technical Details

Source Type
reddit
Subreddit
netsec
Reddit Score
0
Discussion Level
minimal
Content Source
reddit_link_post
Post Type
link
Newsworthiness Assessment
{"score":27,"reasons":["external_link","established_author","very_recent"],"isNewsworthy":true}
Has External Source
true
Trusted Domain
false

Threat ID: 6abb93c7f7a7c541063d3af5

Added to database: 09/29/2026, 10:32:39 UTC

Last enriched: 09/29/2026, 10:32:54 UTC

Last updated: 09/29/2026, 18:07:18 UTC

Views: 16

Community Reviews

0 reviews

Crowdsource mitigation strategies, share intel context, and vote on the most helpful responses. Sign in to add your voice and help keep defenders ahead.

Sort by
Loading community insights…

Want to contribute mitigation steps or threat intel context? Sign in or create an account to join the community discussion.

Actions

PRO

Updates to AI analysis require Pro Console access. Upgrade inside Console → Billing.

Please log in to the Console to use AI analysis features.

Need more coverage?

Upgrade to Pro Console for AI refresh and higher limits.

For incident response and remediation, OffSeq services can help resolve threats faster.

Latest Threats

Breach by OffSeqOFFSEQFRIENDS — 25% OFF

Check if your credentials are on the dark web

Instant breach scanning across billions of leaked records. Free tier available.

Scan now
OffSeq TrainingCredly Certified

Lead Pen Test Professional

Technical5-day eLearningPECB Accredited
View courses