Skip to content

Quick start ​

This guide takes you from zero to a working LibreChat UI backed by GPT-4o-mini, with Palena provisioning the entire backend.

Prerequisites ​

1. Create the working namespace ​

bash
kubectl create namespace palena

2. Create the required Secrets ​

Palena never generates API keys or master passwords on your behalf. Create them up front:

bash
kubectl -n palena create secret generic litellm-keys \
  --from-literal=master=sk-palena-master-$(openssl rand -hex 16) \
  --from-literal=salt=$(openssl rand -hex 32)

kubectl -n palena create secret generic openai-api-key \
  --from-literal=apiKey=sk-your-openai-key

The first Secret holds LiteLLM's master key and encryption salt. The second holds your provider API key — swap in whatever provider you want to expose via the model.

3. Apply a PalenaGateway ​

yaml
apiVersion: operator.palena.ai/v1alpha1
kind: PalenaGateway
metadata:
  name: production
  namespace: palena
spec:
  database:
    managed:
      instances: 1
      storageSize: 10Gi
  redis:
    managed:
      replicas: 1
      storageSize: 5Gi
  gateway:
    replicas: 1
    masterKey:
      name: litellm-keys
      key: master
    saltKey:
      name: litellm-keys
      key: salt
bash
kubectl apply -f gateway.yaml

Watch it come up:

bash
kubectl -n palena get palenagateway production -w

Once PHASE is Running, the gateway is ready. This usually takes 1–2 minutes.

4. Publish a model ​

yaml
apiVersion: operator.palena.ai/v1alpha1
kind: PalenaModel
metadata:
  name: gpt-4o-mini
  namespace: palena
spec:
  gatewayRefs:
    - name: production
  modelName: gpt-4o-mini
  litellmParams:
    model: openai/gpt-4o-mini
    apiKeySecretRef:
      name: openai-api-key
      key: apiKey
bash
kubectl apply -f model.yaml

Verify the model synced to LiteLLM:

bash
kubectl -n palena get palenamodel gpt-4o-mini \
  -o jsonpath='{.status.syncedGateways[0]}{"\n"}'
# {"gatewayName":"production","synced":true,"litellmModelRef":"gpt-4o-mini-production", ...}

5. Add a chat UI (optional) ​

yaml
apiVersion: operator.palena.ai/v1alpha1
kind: PalenaUI
metadata:
  name: chat
  namespace: palena
spec:
  gatewayRef:
    name: production
  replicas: 1
  mongodb:
    managed:
      storageSize: 10Gi
  meiliSearch:
    enabled: true
  ingress:
    enabled: true
    host: chat.example.com
    className: nginx
    tls:
      enabled: true
      secretName: chat-tls
bash
kubectl apply -f ui.yaml

LibreChat, MongoDB, and MeiliSearch come up in sequence. Once PalenaUI.status.ready = true, visit your Ingress host and sign in.

If you don't have an Ingress controller ready, skip the ingress block and port-forward instead:

bash
kubectl -n palena port-forward svc/chat-palena-librechat 3080
# open http://localhost:3080

6. Watch everything come together ​

bash
kubectl -n palena get \
  palenagateway,palenamodel,palenaui

Tearing it down ​

Delete the higher-level CRs first so finalizers can clean up owned state:

bash
kubectl -n palena delete palenaui chat
kubectl -n palena delete palenamodel gpt-4o-mini
kubectl -n palena delete palenagateway production

Palena will wait for each child CR to finish cleanup before removing the parent. The operator also handles upstream CRs (LiteLLMInstance, CNPG Cluster) as part of the finalizer logic — you don't need to delete them manually.

Where to go from here ​

Released under the Apache 2.0 License. "Palena" is a trademark of bitkaio LLC.