Quick start
This guide takes you from zero to a working LibreChat UI backed by GPT-4o-mini, with Palena provisioning the entire backend.
Prerequisites
- A Kubernetes cluster (≥ 1.24) with at least 2 CPU / 4Gi RAM free.
kubectlconfigured for that cluster.- CloudNativePG and LiteLLM Operator installed.
- The Palena Operator installed.
1. Create the working namespace
kubectl create namespace palena2. Create the required Secrets
Palena never generates API keys or master passwords on your behalf. Create them up front:
kubectl -n palena create secret generic litellm-keys \
--from-literal=master=sk-palena-master-$(openssl rand -hex 16) \
--from-literal=salt=$(openssl rand -hex 32)
kubectl -n palena create secret generic openai-api-key \
--from-literal=apiKey=sk-your-openai-keyThe first Secret holds LiteLLM's master key and encryption salt. The second holds your provider API key — swap in whatever provider you want to expose via the model.
3. Apply a PalenaGateway
apiVersion: operator.palena.ai/v1alpha1
kind: PalenaGateway
metadata:
name: production
namespace: palena
spec:
database:
managed:
instances: 1
storageSize: 10Gi
redis:
managed:
replicas: 1
storageSize: 5Gi
gateway:
replicas: 1
masterKey:
name: litellm-keys
key: master
saltKey:
name: litellm-keys
key: saltkubectl apply -f gateway.yamlWatch it come up:
kubectl -n palena get palenagateway production -wOnce PHASE is Running, the gateway is ready. This usually takes 1–2 minutes.
4. Publish a model
apiVersion: operator.palena.ai/v1alpha1
kind: PalenaModel
metadata:
name: gpt-4o-mini
namespace: palena
spec:
gatewayRefs:
- name: production
modelName: gpt-4o-mini
litellmParams:
model: openai/gpt-4o-mini
apiKeySecretRef:
name: openai-api-key
key: apiKeykubectl apply -f model.yamlVerify the model synced to LiteLLM:
kubectl -n palena get palenamodel gpt-4o-mini \
-o jsonpath='{.status.syncedGateways[0]}{"\n"}'
# {"gatewayName":"production","synced":true,"litellmModelRef":"gpt-4o-mini-production", ...}5. Add a chat UI (optional)
apiVersion: operator.palena.ai/v1alpha1
kind: PalenaUI
metadata:
name: chat
namespace: palena
spec:
gatewayRef:
name: production
replicas: 1
mongodb:
managed:
storageSize: 10Gi
meiliSearch:
enabled: true
ingress:
enabled: true
host: chat.example.com
className: nginx
tls:
enabled: true
secretName: chat-tlskubectl apply -f ui.yamlLibreChat, MongoDB, and MeiliSearch come up in sequence. Once PalenaUI.status.ready = true, visit your Ingress host and sign in.
If you don't have an Ingress controller ready, skip the ingress block and port-forward instead:
kubectl -n palena port-forward svc/chat-palena-librechat 3080
# open http://localhost:30806. Watch everything come together
kubectl -n palena get \
palenagateway,palenamodel,palenauiTearing it down
Delete the higher-level CRs first so finalizers can clean up owned state:
kubectl -n palena delete palenaui chat
kubectl -n palena delete palenamodel gpt-4o-mini
kubectl -n palena delete palenagateway productionPalena will wait for each child CR to finish cleanup before removing the parent. The operator also handles upstream CRs (LiteLLMInstance, CNPG Cluster) as part of the finalizer logic — you don't need to delete them manually.
Where to go from here
- Add
PalenaObservabilityto get Langfuse tracing for every request. - Add
PalenaMCPServerto give LibreChat websearch capabilities. - Browse the sample catalog for ready-to-apply manifests covering more scenarios.