ai machinelearning platformengineering distributedsystem reading/researchpaper

Ray for ML Infrastructure — Ray 2.55.1

1. Ray

Ray is an open-source, unified compute framework for scaling Python and AI workloads from a laptop to clusters of thousands of nodes, originally built at UC Berkeley’s RISELab and commercialized by Anyscale.
It provides the compute layer for parallel processing so teams don’t need to be distributed-systems experts, and it minimizes the complexity of running distributed individual workflows and end-to-end ML workflows. In October 2025, Ray was donated to the Linux Foundation and became a PyTorch Foundation-hosted project, signaling broad enterprise and ecosystem backing.

Architecture and core abstractions

  • A simple @ray.remote decorator pattern turns ordinary Python functions/classes into distributed tasks (stateless) and actors (stateful), letting existing code scale with minimal rewrites.
  • Ray automatically handles orchestration, scheduling, and fault tolerance. the hard parts of distributed systems, behind that simple API.
  • It distributes the task scheduler and metadata store across the cluster (rather than centralizing them like Spark or Dask), enabling millions of tasks per second at millisecond-level latency, with lineage-based fault tolerance for tasks/actors.
  • It deploys on AWS, GCP, Azure, or on-prem, and can run on existing Kubernetes, YARN, or Slurm clusters.

Native libraries (the “Ray AI ecosystem”)

  • Ray Data — distributed, framework-agnostic data loading/transformation for ETL and multimodal preprocessing.
  • Ray Train — distributed model training (PyTorch, TensorFlow, etc.).
  • Ray Tune — distributed hyperparameter tuning/AutoML.
  • Ray Serve — scalable, online model/LLM inference serving.
  • Ray RLlib — distributed reinforcement learning(10 - Reinforcement Learning).
  • Modin, a drop-in pandas replacement using a Ray backend for faster large-dataset computation, is also part of the broader ecosystem.

Notable adopters: OpenAI uses Ray to coordinate the training of ChatGPT and other models, and the framework scales from a laptop to clusters of thousands of GPUs. Ray now powers AI workloads at thousands of companies including Uber, xAI, Cursor, Perplexity, and Coinbase, alongside well-documented case studies from Roblox, Airbnb, eBay, Spotify, and Reddit.

Ray: A Distributed Framework for Emerging AI Applications

2. KubeRay

KubeRay is an open-source Kubernetes operator that simplifies the deployment and management of Ray applications on Kubernetes. It’s the standard bridge between Ray’s compute model and Kubernetes-native platform engineering.

Core building blocks

Three Custom Resource Definitions (CRDs):

  • RayCluster — KubeRay fully manages the lifecycle of a Ray cluster, including creation/deletion, autoscaling, and fault tolerance.
  • RayJob — KubeRay automatically creates a RayCluster and submits a job once the cluster is ready, and can be configured to auto-delete the cluster when the job finishes (good for ephemeral batch/training jobs).
  • RayService — combines a RayCluster with a Ray Serve deployment graph, offering zero-downtime upgrades and high availability for production inference endpoints. A community-maintained RayCronJob CRD also exists for scheduled workloads.

Ecosystem integration: KubeRay integrates with observability tools (Promethues, Grafana, py-spy), queuing/scheduling systems (Volcano, Apache YuniKorn, Kueue), and ingress controllers (NGINX), and supports heterogeneous compute nodes (CPU/GPU mixes) plus multiple Ray versions on the same Kubernetes cluster.

What’s new (KubeRay v1.4): A KubeRay API Server V2 lets platform engineers build user interfaces for data scientists who lack direct Kubernetes API access, plus a Ray Autoscaler V2 and SLI metrics — directly aimed at the platform-engineering pain point of hiding Kubernetes complexity from ML practitioners while still answering SLA questions like cluster-startup time.

3. Anyscale on Azure

Anyscale (founded by Ray’s creators) layers a managed control plane, performance runtime, and enterprise tooling on top of open-source Ray. In Nov 2025 Microsoft and Anyscale announced a co-engineered, first-party Azure integration, which reached public preview on June 2, 2026.

Architecture

  • Control plane: Anyscale hosts this in Azure, handling scheduling, monitoring, job management, and the Anyscale console, accessible via the Azure portal, CLI, or SDK.
  • Data plane: runs inside the customer’s own Azure subscription, on the customer’s AKS cluster — a bring-your-own-cloud (BYOC) model that keeps workload data and compute inside the customer’s tenant.
  • It supports unified workload types — Ray Data for ETL, Ray Train for distributed training, Ray Serve for online inference, and Ray RLlib for reinforcement learning — all on the same cluster.

Azure-native integration

  • Native authentication via Azure Entra ID, with Anyscale-managed Ray workloads running on AKS under the organization’s existing IAM policies and unified governance.
  • Authentication has moved from expiring CLI tokens/API keys to Microsoft Entra service principals and AKS workload identity, issuing short-lived tokens automatically and producing full audit trails via Azure Activity Logs.
  • Sovereign AI controls — data residency guarantees, customer-managed encryption keys, and deployment in sovereign regions (e.g., Azure Government, Azure Germany) — target regulated and public-sector customers.
  • Pricing is pay-as-you-go through Azure service meters (no upfront commitment), billed on CPU, memory, and GPU type, with Anyscale charging for the orchestration/management layer rather than GPU capacity itself.

Performance — the “Anyscale Runtime”

  • A Ray-compatible runtime optimized for higher performance and reliability than open-source Ray, achieving up to 10x faster feature preprocessing and batch image inference, with no code changes required.
  • It improves stability for long-running, large-scale jobs via checkpointing (pause/resume batch processing), mid-epoch resume for training, and dynamic memory management that reduces spilling and out-of-memory errors.

Addressing real operational pain points (per Microsoft/AKS guidance)

  • GPU scarcity: a multi-cluster, multi-region setup lets teams aggregate GPU quota beyond regional limits, automatically reroute workloads during outages, and extend the compute pool to on-prem or other clouds via Azure Arc with AKS.
  • Storage portability: enabling the Blob CSI driver and a workload-identity-authenticated StorageClass lets multiple Ray workers across nodes share data via a ReadWriteMany PersistentVolumeClaim.

4. Platform Engineering use cases

LayerWhat it gives platform teamsTypical use case
Ray (OSS)A single Pythonic compute substrate instead of bespoke distributed-systems code per teamStandardizing how every ML/data team scales batch ETL, training, hyperparameter search, RL, and inference — one mental model instead of N custom frameworks
KubeRayKubernetes-native lifecycle management (CRDs) for Ray clusters, jobs, and servicesBuilding a self-service “ML platform on K8s”: RayJob for ephemeral training/batch runs that clean up after themselves; RayCluster for shared interactive dev environments; RayService for zero-downtime, autoscaling LLM/model-serving endpoints; integrating with existing GitOps (ArgoCD), queueing (Kueue/Volcano/YuniKorn), and observability (Prometheus/Grafana) stacks the platform team already runs
Anyscale on AzureA managed control plane + governed data plane, removing day-2 ops burden while keeping compute inside the org’s Azure tenantEnterprise/regulated environments that want Ray’s flexibility but need centralized governance: Entra-based RBAC and audit trails, sovereign-region deployment, cost chargeback per team/project, multi-region GPU capacity pooling, and a portal-native experience so platform teams don’t have to build their own KubeRay abstraction layer from scratch

Concrete platform-engineering scenarios

  1. Self-service ML/AI compute platform — Platform teams expose RayJob/RayService templates (via Helm, Argo, or an internal portal) so data scientists submit training or inference workloads without touching the Kubernetes API directly. This is explicitly the gap KubeRay v1.4’s API Server V2 is designed to close — letting platform engineers build UIs for users without direct K8s API server access.
  2. Unified batch + online inference fabric — Use RayJob/RayCluster for batch inference and large-scale data processing pipelines, and RayService for production-grade, autoscaling LLM serving with an OpenAI-compatible API, all on the same cluster definition.
  3. Cost and capacity governance — On Anyscale on Azure, platform teams get granular cost tracking per project, team, or user, with budget alerts and auto-shutdown policies to prevent runaway GPU spend, plus dashboards for resource utilization and spend breakdowns.
  4. GPU capacity orchestration across regions — Distributing Ray clusters across multiple AKS instances in different Azure regions lets platform teams aggregate scarce GPU SKUs beyond a single region’s quota and fail over automatically during outages.
  5. Regulated/sovereign workloads — For government, defense, or finance customers, Anyscale on Azure’s sovereign-region and customer-managed-key support lets platform teams offer a managed Ray experience without violating data-residency requirements.
  6. Hybrid/on-prem extension — Azure Arc with AKS can extend the Ray/KubeRay compute pool to on-premises systems or other clouds, useful for platform teams managing a mixed estate rather than a pure public-cloud footprint.
  7. MLOps lifecycle backbone — KubeRay can be used to manage the end-to-end lifecycle of ML/LLM models — experimentation, data processing, training, and serving — built as a custom AI/ML platform layered on top of Kubernetes for reliability.

5. How the three pieces relate

  • Ray = the distributed-compute library/runtime your application code targets (the “what”).
  • KubeRay = the Kubernetes operator that turns Ray clusters/jobs/services into declarative, GitOps-friendly K8s resources (the “how,” self-managed, free, open source).
  • Anyscale on Azure = a managed product built on Ray/KubeRay concepts, adding a hosted control plane, a faster proprietary runtime, enterprise governance (Entra ID, sovereign regions, billing), and removing most day-2 operational burden from the platform team — the trade-off being a managed-service relationship (SLAs, vendor dependency) instead of full self-hosted control.

A typical adoption path for a platform engineering org: start with Ray locally to validate the workload, move to KubeRay on an existing AKS/EKS/GKE cluster for self-managed production use, and consider Anyscale on Azure (or Anyscale’s other cloud offerings) when the operational overhead of running KubeRay at scale, multi-region GPU contention, or compliance/governance requirements outgrow what an internal platform team wants to build and maintain itself.


Q&A