Skip to content

CRD design ​

This page explains the shape of the Palena CRDs — their relationships, default values, and enforcement rules. For a field-by-field reference, see the CRD Reference.

Reference graph ​

  • PalenaModel → PalenaGateway is many-to-many. A single PalenaModel can be published to multiple gateways via gatewayRefs. The controller creates one LiteLLMModel CR per (model, gateway) pair.
  • PalenaObservability → PalenaGateway is one-to-one. One observability CR per gateway.
  • PalenaUI → PalenaGateway is one-to-one. One UI binds to exactly one gateway.
  • PalenaUI → PalenaMCPServer is one-to-many. A UI can register any number of MCP servers for tool use.
  • PalenaMCPServer is standalone. It does not reference a gateway in v1alpha1.

Deletion order ​

Finalizers enforce a strict teardown order:

PalenaUI      ─┐
PalenaModel   ─┤
PalenaObservability ─┼──► PalenaGateway (blocked until all children gone)
               ─┘

When you kubectl delete palenagateway production, the gateway's finalizer blocks removal until:

  1. All PalenaModel CRs referencing this gateway are deleted (or their gatewayRefs updated to exclude it).
  2. All PalenaUI CRs pointing at this gateway are deleted.
  3. All PalenaObservability CRs tied to this gateway are deleted, and the Langfuse callback has been removed from the LiteLLMInstance.

Only then does the gateway finalizer run its own cleanup (delete CNPG Cluster, Redis resources, LiteLLMInstance) and remove itself.

PalenaMCPServer is not part of this chain — it's standalone. When deleted, its owner references alone are enough to GC the managed resources.

Per-CRD shape ​

PalenaGateway ​

The platform foundation. Deploys PostgreSQL, Redis, and the LLM gateway (LiteLLM).

yaml
spec:
  database:         # exactly one of managed | external
    managed:
      instances: 1
      storageSize: 10Gi
  redis:            # exactly one of managed | external
    managed:
      replicas: 1
      storageSize: 5Gi
  gateway:
    replicas: 1
    masterKey: { name: litellm-keys, key: master }   # required
    saltKey:   { name: litellm-keys, key: salt }     # recommended
    configSync:
      enabled: true
      mode: bidirectional        # enum: bidirectional | gitops-only | ui-only
      interval: 30s
      unmanagedResourcePolicy: preserve   # enum: preserve | prune | adopt
    routerSettings:
      routingStrategy: simple-shuffle
      numRetries: 2
      timeout: 60
    ingress:
      enabled: true
      host: llm.example.com
      className: nginx
  security:
    networkPolicies:
      enabled: true
  platform: auto    # enum: auto | kubernetes | openshift

Status phases: Pending → Provisioning → Running → Degraded / Error.

PalenaModel ​

Defines a model available through one or more gateways. Minimum spec:

yaml
spec:
  gatewayRefs:
    - name: production
  modelName: gpt-4o-mini           # logical name exposed by LiteLLM
  litellmParams:
    model: openai/gpt-4o-mini      # upstream LiteLLM model identifier
    apiKeySecretRef:
      name: openai-api-key
      key: apiKey

The controller creates one LiteLLMModel CR per gatewayRef, named <modelName>-<gatewayName>. Removing a name from gatewayRefs triggers orphan cleanup — the controller deletes any LiteLLMModel CRs it owned that no longer match.

PalenaUI ​

Chat interface. References a PalenaGateway and optionally PalenaMCPServers.

yaml
spec:
  gatewayRef:
    name: production
  mcpServerRefs:
    - name: websearch
  replicas: 1
  mongodb:
    managed:
      storageSize: 10Gi
  meiliSearch:
    enabled: true
  branding:
    appTitle: Acme Chat
  auth:
    disableSignup: true
    oidc:
      enabled: true
      issuer: https://auth.example.com/realms/acme
      clientCredentials:
        name: oidc-creds
        key: clientSecret
  ingress:
    enabled: true
    host: chat.example.com
    className: nginx

The operator auto-generates a librechat.yaml ConfigMap and attaches its SHA-256 as a pod annotation — any config change triggers a rolling restart.

PalenaObservability ​

Langfuse tracing tied to a gateway.

yaml
spec:
  gatewayRef:
    name: production
  langfuse:
    web:
      replicas: 1
    worker:
      replicas: 1
    clickhouse:
      managed:
        replicas: 1
        storageSize: 50Gi
    blobStorage:
      provider: s3      # enum: s3 | azure | gcs
      s3:
        region: us-east-1
        bucket: palena-langfuse-traces
        credentials:
          name: s3-creds
          keys:
            accessKeyId: AWS_ACCESS_KEY_ID
            secretAccessKey: AWS_SECRET_ACCESS_KEY
    auth:
      initUser:
        email: admin@example.com
        password:
          name: langfuse-init
          key: password

After the LangfuseInstance is ready, the controller seeds an admin user, extracts the generated API keys, and patches the gateway's LiteLLMInstance to add the Langfuse callback.

PalenaMCPServer ​

MCP server deployment. Only type: websearch in v1alpha1.

yaml
spec:
  type: websearch
  replicas: 1
  websearch:
    searxng:
      replicas: 1
      engines: [duckduckgo, bing, wikipedia]
    scraper:
      chromiumEnabled: true
      chromiumReplicas: 1
      maxConcurrency: 10
    presidio:
      enabled: true
      mode: redact          # enum: audit | redact | block
    reranker:
      provider: flashrank   # enum: flashrank | kserve | none

Each sub-component (SearXNG, Chromium, Presidio, FlashRank) becomes an independent Deployment + Service. The MCP server itself loads a generated config ConfigMap pointing at all enabled sidecars.

Versioning ​

The Palena API is currently v1alpha1. Breaking changes may occur between minor versions until the API graduates to v1beta1.

When new fields are added:

  1. They will be optional with a sensible default.
  2. Deprecations are announced at least one minor release before removal.
  3. Release notes call out any migration steps.

Pin to a specific bundle/chart tag in production and test upgrades in staging first.

Released under the Apache 2.0 License. "Palena" is a trademark of bitkaio LLC.