Cloud deployment - Poolside

Supported environments

Amazon EKS
Deploy model inference with Helm on Amazon EKS, using IRSA for object storage and an Application Load Balancer for ingress.

Red Hat OpenShift
Deploy model inference with Helm on your OpenShift cluster.

Upstream Kubernetes
Deploy model inference with Helm on your self-managed Kubernetes cluster, such as RKE2 or Charmed Kubernetes.

Architecture

Cloud deployment includes:

You are responsible for sending requests to the inference endpoints and for any authentication or routing in front of them. To expose a shared endpoint with centralized access controls, routing, virtual keys, budgets, and spend tracking, see Use LiteLLM with Poolside model inference.

Operational considerations