CRD design
This page explains the shape of the Palena CRDs — their relationships, default values, and enforcement rules. For a field-by-field reference, see the CRD Reference.
Reference graph
- PalenaModel → PalenaGateway is many-to-many. A single PalenaModel can be published to multiple gateways via
gatewayRefs. The controller creates oneLiteLLMModelCR per (model, gateway) pair. - PalenaObservability → PalenaGateway is one-to-one. One observability CR per gateway.
- PalenaUI → PalenaGateway is one-to-one. One UI binds to exactly one gateway.
- PalenaUI → PalenaMCPServer is one-to-many. A UI can register any number of MCP servers for tool use.
- PalenaMCPServer is standalone. It does not reference a gateway in v1alpha1.
Deletion order
Finalizers enforce a strict teardown order:
PalenaUI ─┐
PalenaModel ─┤
PalenaObservability ─┼──► PalenaGateway (blocked until all children gone)
─┘When you kubectl delete palenagateway production, the gateway's finalizer blocks removal until:
- All
PalenaModelCRs referencing this gateway are deleted (or theirgatewayRefsupdated to exclude it). - All
PalenaUICRs pointing at this gateway are deleted. - All
PalenaObservabilityCRs tied to this gateway are deleted, and the Langfuse callback has been removed from theLiteLLMInstance.
Only then does the gateway finalizer run its own cleanup (delete CNPG Cluster, Redis resources, LiteLLMInstance) and remove itself.
PalenaMCPServer is not part of this chain — it's standalone. When deleted, its owner references alone are enough to GC the managed resources.
Per-CRD shape
PalenaGateway
The platform foundation. Deploys PostgreSQL, Redis, and the LLM gateway (LiteLLM).
spec:
database: # exactly one of managed | external
managed:
instances: 1
storageSize: 10Gi
redis: # exactly one of managed | external
managed:
replicas: 1
storageSize: 5Gi
gateway:
replicas: 1
masterKey: { name: litellm-keys, key: master } # required
saltKey: { name: litellm-keys, key: salt } # recommended
configSync:
enabled: true
mode: bidirectional # enum: bidirectional | gitops-only | ui-only
interval: 30s
unmanagedResourcePolicy: preserve # enum: preserve | prune | adopt
routerSettings:
routingStrategy: simple-shuffle
numRetries: 2
timeout: 60
ingress:
enabled: true
host: llm.example.com
className: nginx
security:
networkPolicies:
enabled: true
platform: auto # enum: auto | kubernetes | openshiftStatus phases: Pending → Provisioning → Running → Degraded / Error.
PalenaModel
Defines a model available through one or more gateways. Minimum spec:
spec:
gatewayRefs:
- name: production
modelName: gpt-4o-mini # logical name exposed by LiteLLM
litellmParams:
model: openai/gpt-4o-mini # upstream LiteLLM model identifier
apiKeySecretRef:
name: openai-api-key
key: apiKeyThe controller creates one LiteLLMModel CR per gatewayRef, named <modelName>-<gatewayName>. Removing a name from gatewayRefs triggers orphan cleanup — the controller deletes any LiteLLMModel CRs it owned that no longer match.
PalenaUI
Chat interface. References a PalenaGateway and optionally PalenaMCPServers.
spec:
gatewayRef:
name: production
mcpServerRefs:
- name: websearch
replicas: 1
mongodb:
managed:
storageSize: 10Gi
meiliSearch:
enabled: true
branding:
appTitle: Acme Chat
auth:
disableSignup: true
oidc:
enabled: true
issuer: https://auth.example.com/realms/acme
clientCredentials:
name: oidc-creds
key: clientSecret
ingress:
enabled: true
host: chat.example.com
className: nginxThe operator auto-generates a librechat.yaml ConfigMap and attaches its SHA-256 as a pod annotation — any config change triggers a rolling restart.
PalenaObservability
Langfuse tracing tied to a gateway.
spec:
gatewayRef:
name: production
langfuse:
web:
replicas: 1
worker:
replicas: 1
clickhouse:
managed:
replicas: 1
storageSize: 50Gi
blobStorage:
provider: s3 # enum: s3 | azure | gcs
s3:
region: us-east-1
bucket: palena-langfuse-traces
credentials:
name: s3-creds
keys:
accessKeyId: AWS_ACCESS_KEY_ID
secretAccessKey: AWS_SECRET_ACCESS_KEY
auth:
initUser:
email: admin@example.com
password:
name: langfuse-init
key: passwordAfter the LangfuseInstance is ready, the controller seeds an admin user, extracts the generated API keys, and patches the gateway's LiteLLMInstance to add the Langfuse callback.
PalenaMCPServer
MCP server deployment. Only type: websearch in v1alpha1.
spec:
type: websearch
replicas: 1
websearch:
searxng:
replicas: 1
engines: [duckduckgo, bing, wikipedia]
scraper:
chromiumEnabled: true
chromiumReplicas: 1
maxConcurrency: 10
presidio:
enabled: true
mode: redact # enum: audit | redact | block
reranker:
provider: flashrank # enum: flashrank | kserve | noneEach sub-component (SearXNG, Chromium, Presidio, FlashRank) becomes an independent Deployment + Service. The MCP server itself loads a generated config ConfigMap pointing at all enabled sidecars.
Versioning
The Palena API is currently v1alpha1. Breaking changes may occur between minor versions until the API graduates to v1beta1.
When new fields are added:
- They will be optional with a sensible default.
- Deprecations are announced at least one minor release before removal.
- Release notes call out any migration steps.
Pin to a specific bundle/chart tag in production and test upgrades in staging first.