Install on Amazon EKS - Poolside

Follow these steps to deploy Poolside model inference on your Amazon EKS cluster and serve models through an OpenAI-compatible API. For an overview of this deployment approach and architecture, see Amazon EKS deployment. This guide deploys the Poolside inference chart. Each model becomes its own Deployment, Service, and Ingress, reachable at its own hostname through a shared Application Load Balancer.

Prerequisites

Poolside distributes the Helm deployment bundle as a .tar.gz archive. Extract it before you start:

tar -xzf <bundle-name>.tar.gz
cd <bundle-name>

Confirm that you are working from the root of the extracted bundle. The bundle root contains the following directories:

./scripts/
./containers/
./charts/

Required AWS infrastructure

You provision the following AWS foundation before you deploy the chart. For a turnkey foundation that provisions all of it, apply the Terraform reference architecture in the poolsideai/reference_architectures repository, or reproduce the same architecture in your own infrastructure-as-code. For the architecture diagram and design decisions, see Reference architecture.

Workstation tools

Install the following tools on the host you use to run the deployment:

Disk space: Stage the deployment from a host with tens of GB free. The extracted bundle is roughly 20 GB.

Step 1: Create the namespace

The inference stack runs in a single namespace:

kubectl create namespace poolside-models

Step 2: Upload the container images to Amazon ECR

Copy the images into Amazon ECR. For each image name from the bundle:

for image_name in $(find ./containers -name "*.tar" -type f -exec basename {} .tar \; | sed 's/__.*//' | sort -u); do
  aws ecr describe-repositories --repository-names "$image_name" --region <aws-region> >/dev/null 2>&1 \
    || aws ecr create-repository --repository-name "$image_name" --region <aws-region>
done

Authenticate skopeo to your ECR registry:

aws ecr get-login-password --region <aws-region> \
  | skopeo login --username AWS --password-stdin <account-id>.dkr.ecr.<aws-region>.amazonaws.com

Upload the images with the provided script:

chmod +x ./scripts/upload_images.sh
./scripts/upload_images.sh <account-id>.dkr.ecr.<aws-region>.amazonaws.com

Step 3: Upload model checkpoints to S3

The model servers download their checkpoints from S3 on pod startup. Upload the checkpoints to your S3 bucket:

aws s3 cp ./checkpoints s3://<bucket-name>/checkpoints --recursive --region <aws-region>

Step 4: Create the IRSA role

Save the permissions and trust policy to respective JSON files. Create the policy and role:

aws iam create-policy \
  --policy-name inference-pod-policy \
  --policy-document file://inference-pod-policy.json

aws iam create-role \
  --role-name inference-pod-role \
  --assume-role-policy-document file://inference-pod-trust.json

aws iam attach-role-policy \
  --role-name inference-pod-role \
  --policy-arn arn:aws:iam::<account-id>:policy/inference-pod-policy

Step 5: Create the API key secret (recommended for internet-facing)

Create a secret containing the API key:

kubectl create secret generic vllm-auth \
  --from-literal=VLLM_API_KEY=<vllm-api-key> \
  -n poolside-models

Step 6: Configure the values file

Create an inference_values.yaml file in the bundle root:

fullnameOverride: inference

image:
  registry: <account-id>.dkr.ecr.<aws-region>.amazonaws.com

serviceAccount:
  create: true
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::<account-id>:role/inference-pod-role

authentication:
  secretName: vllm-auth

ingress:
  enabled: true
  className: alb
  annotations:
    alb.ingress.kubernetes.io/scheme: internet-facing

models:
  laguna-m:
    model: s3://<bucket-name>/checkpoints/laguna-m
    modelName: Lagunam
    modelType: agent
    gpus: 4
  laguna-xs:
    model: s3://<bucket-name>/checkpoints/laguna-xs
    modelName: Lagunaxs
    modelType: agent
    gpus: 1

Step 7: Install the chart

Install the inference chart:

helm install inference ./charts/inference \
  --namespace poolside-models \
  -f ./inference_values.yaml

Step 8: Verify the deployment

Check model pods:

kubectl get pods -n poolside-models

Step 9: Call the inference API

Each model serves an API at its own hostname.

Send a request to a model:

curl https://<model-hostname>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<served-model-name>",
    "messages": [{"role": "user", "content": "Write a function that reverses a string."}]
  }'

TLS

The load balancer terminates TLS with the ACM certificate and serves models over HTTPS.