## Prerequisites

Poolside distributes the Helm deployment bundle as a `.tar.gz` archive. Extract it before you start:

```bash
tar -xzf <bundle-name>.tar.gz
cd <bundle-name>
```

Confirm that you are working from the root of the extracted bundle. The bundle root contains the following directories:

```bash
./scripts/
./containers/
./charts/
./binaries/
```

**Cluster requirements**

- Kubernetes 1.29 or later
- GPU nodes with enough GPUs for the models you deploy
- NVIDIA GPU Operator 26.3.0, with NVIDIA driver 580.126.20 and NVIDIA Container Toolkit 1.19.0
- An ingress controller that can route HTTP and HTTPS traffic to the cluster
- A DNS hostname for each model you deploy, resolving to the ingress endpoint. Kubernetes Ingress objects do not accept bare IP addresses; use a DNS name or `/etc/hosts` entries.
- An S3-compatible object storage service such as Amazon S3, SeaweedFS, MinIO, or NooBaa
- A container registry that every cluster node can access

**Workstation tools**
Install the following tools on the host you use to run the deployment:

- `helm` `3.12` or later
- `kubectl`
- `skopeo`
- `aws` CLI (to upload checkpoints to S3-compatible object storage)
- `jq` (to parse JSON responses from the inference API)
- `tar` (to extract the deployment bundle)
- `curl` (to call the inference API)
- `openssl` (optional, to generate a TLS certificate for the inference endpoint)

**Minimum resource requirements**
Ensure that your cluster has enough GPUs for the models you deploy. If you have questions about the required specs, contact Poolside support.

## Step 1: Create the namespace

The inference stack runs in a single namespace:

```bash
kubectl create namespace poolside-models
```

## Step 2: Upload container images

Copy the bundled images into your registry. Log in to your target registry using `docker login` or `podman login` before running any upload commands.

Authenticate skopeo against your target registry:

```bash
skopeo login <registry-host> --username <username> --password <password>
```

Upload the images with the provided script:

```bash
chmod +x ./scripts/upload_images.sh
./scripts/upload_images.sh <registry-host>
```

If your registry does not use TLS:

```bash
./scripts/upload_images.sh <registry-host>:5000 --force-insecure-dest
```

If your registry requires authentication, create an image pull secret in `poolside-models`:

```bash
kubectl create secret docker-registry poolside-registry-secret \
  --docker-server=<registry-host> \
  --docker-username=<registry-user> \
  --docker-password=<registry-password> \
  -n poolside-models
```

## Step 3: Upload model checkpoints

The inference stack downloads model weights from your S3 bucket on pod startup, so the checkpoints must be in place before you deploy the chart. Poolside provides the checkpoint files separately from the deployment bundle. Confirm the local path and the destination prefix with your Poolside contact. Uploading checkpoints is time consuming. Start it now and continue with the remaining steps in parallel. Poolside provides model checkpoints as `.tar` archives. The inference chart does not extract archives at pod startup. It syncs unpacked checkpoint files (`*.safetensors`, `*.json`, and the tokenizer files) from the S3 prefix you set in `models.<key>.model`, so extract each archive’s contents into its own directory before uploading.

```bash
tar -xf ./checkpoints/<checkpoint-archive>.tar
rm -v ./checkpoints/<checkpoint-archive>.tar
```

Confirm the files sit at the root of the directory, not under a subfolder:

```bash
ls ./checkpoints/<model-key>
# config.yaml  generation_config.json  model.safetensors  tokenizer/
```

Create the bucket if it does not already exist:

```bash
aws s3 mb s3://<bucket-name> --region <aws-region>
```

For a non-AWS S3 endpoint (MinIO, NooBaa, SeaweedFS), add `--endpoint-url https://<s3-endpoint>`.

Then upload the checkpoints to the bucket:

```bash
aws s3 cp ./checkpoints s3://<bucket-name>/checkpoints --recursive --region <aws-region>
```

For a non-AWS S3 endpoint (MinIO, NooBaa, SeaweedFS), add `--endpoint-url`:

```bash
aws s3 cp ./checkpoints s3://<bucket-name>/checkpoints \
  --recursive \
  --endpoint-url https://<s3-endpoint> \
  --region <aws-region>
```

Checkpoints are typically tens of GiB per model. For faster throughput, or for backends sensitive to upload concurrency such as NooBaa or SeaweedFS, run the upload from a host inside the cluster and tune `aws configure set default.s3.max_concurrent_requests` and `default.s3.multipart_chunksize`.

## Step 4: Create the S3 credentials secret

The model servers read checkpoints from S3 using credentials in a Kubernetes secret. Create it in `poolside-models`:

```bash
kubectl create secret generic aws-credentials \
  --from-literal=AWS_ACCESS_KEY_ID=<access-key-id> \
  --from-literal=AWS_SECRET_ACCESS_KEY=<secret-access-key> \
  -n poolside-models
```

**API key authentication (optional)**
To require an API key on the vLLM inference servers, create a secret containing the key in `poolside-models`:

```bash
kubectl create secret generic vllm-auth \
  --from-literal=VLLM_API_KEY=<vllm-api-key> \
  -n poolside-models
```

Creating the secret does not enable API key authentication by itself. Set `authentication.secretName` to `vllm-auth` in your values file.

## Step 5: Configure the inference values file

Create an `inference_values.yaml` file in the bundle root:

```bash
cp ./charts/inference/values.yaml ./inference_values.yaml
```

Set the fields that apply to your environment. The example below deploys two models and exposes each model through its own ingress:

Example: `inference_values.yaml`

```yaml
image:
  registry: "<registry-host>"
  name: "atlas"
  tag: "<atlas-tag>"
imagePullSecret: "poolside-registry-secret"
podSecurityContext:
  runAsNonRoot: true
  runAsUser: 10003
  seccompProfile:
    type: RuntimeDefault
s3:
  secretName: "aws-credentials"
  caBundle: ""
authentication:
  secretName: ""
ingress:
  enabled: true
  className: "nginx"
models:
  laguna-m:
    model: s3://<bucket-name>/checkpoints/laguna-m
    modelName: Lagunam
    modelType: agent
    gpus: 4
    ingressHost: "<laguna-m-hostname>"
  laguna-xs:
    model: s3://<bucket-name>/checkpoints/laguna-xs
    modelName: Lagunaxs
    modelType: agent
    gpus: 1
    ingressHost: "<laguna-xs-hostname>"
```

The checkpoint paths in `models.<key>.model` and the image registry must exactly match the locations you uploaded from the deployment bundle. Set each model’s `gpus` to a value that meets its minimum GPU memory for your GPU type. For the per-model minimums, see Supported configurations.

Each model is exposed at its own hostname through a separate Ingress named `inference-<model-key>`. Give every model a unique `ingressHost`. The Ingress routes the hostname’s root path directly to that model’s vLLM service, so clients reach the OpenAI-compatible API at `http://<model-hostname>/v1`.

**Non-AWS S3 endpoints**
If your object storage is not AWS S3, point the model servers at the endpoint and region:

```yaml
extraEnv:
  AWS_REGION: "<aws-region>"
  AWS_ENDPOINT_URL_S3: "https://<s3-endpoint>"
```

When you use SeaweedFS as the S3 backend, set the AWS CLI to the classic transfer client:

```yaml
awsCliConfig:
  default.s3.preferred_transfer_client: "classic"
```

When you use NooBaa or another S3 backend with limited concurrency, throttle downloads:

```yaml
awsCliConfig:
  default.s3.max_concurrent_requests: "2"
  default.s3.max_queue_size: "1000"
  default.s3.multipart_chunksize: "64MB"
```

## Step 6: Install the inference chart

Install the `inference` chart into `poolside-models`:

```bash
helm install inference ./charts/inference \
  --namespace poolside-models \
  -f ./inference_values.yaml
```

If your S3 backend uses a private CA, include the CA bundle at install time:

```bash
helm install inference ./charts/inference \
  --namespace poolside-models \
  -f ./inference_values.yaml \
  --set-file s3.caBundle=<path-to-s3-ca.crt>
```

## Step 7: Verify the deployment

Check that the model pods are running. The only pods in the namespace are the per-model servers:

```bash
kubectl get pods -n poolside-models
```

Each model server takes time to become ready on first start because it downloads its checkpoint from S3. Watch a model’s logs to track progress:

```bash
kubectl logs -f -n poolside-models deploy/inference-<model-key>
```

Confirm an ingress was created for each model:

```bash
kubectl get ingress -n poolside-models
```

List the served models on a model’s endpoint to confirm routing works:

```bash
curl -s http://<model-hostname>/v1/models
```

## Step 8: Call the inference API

Each model serves the OpenAI-compatible API directly at its own hostname. The base URL has the form:

```txt
http://<model-hostname>/v1
```

Append the OpenAI-compatible route to the base URL, such as `/chat/completions` or `/completions`. The commands below use three placeholders. Fill them from the `inference_values.yaml` you wrote:

| Placeholder | Source in `inference_values.yaml` | Example |
| --- | --- | --- |
| `<model-hostname>` | `models.<model-key>.ingressHost` | `laguna-m.poolside.local` |
| `<model-key>` | a key under `models` | `laguna-m` |
| `<served-model-name>` | `models.<model-key>.modelName` | `Laguna` |

Send a chat completion request:

```bash
curl http://<model-hostname>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<served-model-name>",
    "messages": [{"role": "user", "content": "Write a function that reverses a string."}]
  }'
```

For example, to call the `laguna-m` model served as `Laguna`:

```bash
curl http://laguna-m.poolside.local/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Laguna",
    "messages": [{"role": "user", "content": "Write a function that reverses a string."}]
  }'
```

If you set `authentication.secretName`, include the key as a bearer token:

```bash
curl http://<model-hostname>/v1/chat/completions \
  -H "Authorization: Bearer <vllm-api-key>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<served-model-name>",
    "messages": [{"role": "user", "content": "Write a function that reverses a string."}]
  }'
```

## TLS

To serve the inference endpoints over HTTPS, add a `tls` block to `ingress`:

```yaml
ingress:
  enabled: true
  className: "nginx"
  tls:
    - hosts:
        - "<laguna-m-hostname>"
      secretName: "<laguna-m-tls-secret>"
    - hosts:
        - "<laguna-xs-hostname>"
      secretName: "<laguna-xs-tls-secret>"
```

Create each referenced secret with `kubectl create secret tls`, or use `cert-manager` to provision it. Clients then reach each model at `https://<model-hostname>/v1`.

## Troubleshooting

- If pods stay in `Init` or restart in a loop, check the init container logs with `kubectl logs -n poolside-models <pod-name> -c <init-container>`. A stale or misspelled checkpoint path syncs nothing and the pod never starts.
- If model pods never become ready, confirm the S3 prefix holds the correct checkpoint files and extract any `.tar` archives as necessary.
- Verify that `imagePullSecret` references the correct secret if model servers fail to pull images.
