Skip to content

PalenaModel ​

A PalenaModel publishes one logical model to one or more PalenaGateways. It translates into one LiteLLMModel CR per referenced gateway.

Short name: pmAPI group: operator.palena.ai/v1alpha1Scope: Namespaced

spec ​

FieldTypeRequiredDescription
gatewayRefs[]LocalObjectReferenceYes (≥1)PalenaGateway names (same namespace).
modelNamestringYesPublic model name exposed by the gateway.
litellmParamsLiteLLMModelParamsYesUpstream LiteLLM model parameters.
modelInfoModelInfoNoOptional token limits and cost metadata.
dataClassificationMaxLevelint (0–4)NoUpper bound for routing policies.
allowedTeams[]stringNoRestrict to named LiteLLM teams.

LiteLLMModelParams ​

FieldTypeRequiredDescription
modelstringYesUpstream model identifier (e.g. openai/gpt-4o-mini).
apiBasestringNoCustom API endpoint for self-hosted models.
apiKeySecretRefSecretKeyRefNoProvider API key secret.
rpmintNoRequests/minute rate limit.
tpmintNoTokens/minute rate limit.
timeoutintNoPer-request timeout (seconds).
streamTimeoutintNoStreaming timeout (seconds).
maxRetriesintNoUpstream retry count.

ModelInfo ​

FieldTypeDescription
maxTokensintTotal context window.
inputCostPerTokenfloat64Input token cost (dollars per token).
outputCostPerTokenfloat64Output token cost (dollars per token).

status ​

FieldTypeDescription
syncedGateways[].gatewayNamestringReferenced gateway.
syncedGateways[].syncedboolSuccessful sync to LiteLLM.
syncedGateways[].litellmModelRefstringName of the generated LiteLLMModel.
syncedGateways[].lastSyncTimestringRFC3339 timestamp.
syncedGateways[].errorstringMost recent sync error, if any.
conditions[]metav1.ConditionStandard conditions (Synced, Ready).

Model, Age.

Example ​

yaml
apiVersion: operator.palena.ai/v1alpha1
kind: PalenaModel
metadata:
  name: gpt-4o-mini
  namespace: palena
spec:
  gatewayRefs:
    - name: production
    - name: staging
  modelName: gpt-4o-mini
  litellmParams:
    model: openai/gpt-4o-mini
    apiKeySecretRef:
      name: openai-keys
      key: apiKey
    rpm: 10000
    tpm: 2000000
  modelInfo:
    maxTokens: 128000
    inputCostPerToken: 0.00000015
    outputCostPerToken: 0.0000006
  dataClassificationMaxLevel: 2

Multi-gateway fan-out ​

When gatewayRefs lists more than one gateway, the controller creates one LiteLLMModel CR per gateway, each named <modelName>-<gatewayName>. Removing a gateway from the list cascades into deletion of the orphaned LiteLLMModel.

Released under the Apache 2.0 License. "Palena" is a trademark of bitkaio LLC.