PalenaModel
A PalenaModel publishes one logical model to one or more PalenaGateways. It translates into one LiteLLMModel CR per referenced gateway.
Short name: pmAPI group: operator.palena.ai/v1alpha1Scope: Namespaced
spec
| Field | Type | Required | Description |
|---|---|---|---|
gatewayRefs | []LocalObjectReference | Yes (≥1) | PalenaGateway names (same namespace). |
modelName | string | Yes | Public model name exposed by the gateway. |
litellmParams | LiteLLMModelParams | Yes | Upstream LiteLLM model parameters. |
modelInfo | ModelInfo | No | Optional token limits and cost metadata. |
dataClassificationMaxLevel | int (0–4) | No | Upper bound for routing policies. |
allowedTeams | []string | No | Restrict to named LiteLLM teams. |
LiteLLMModelParams
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Upstream model identifier (e.g. openai/gpt-4o-mini). |
apiBase | string | No | Custom API endpoint for self-hosted models. |
apiKeySecretRef | SecretKeyRef | No | Provider API key secret. |
rpm | int | No | Requests/minute rate limit. |
tpm | int | No | Tokens/minute rate limit. |
timeout | int | No | Per-request timeout (seconds). |
streamTimeout | int | No | Streaming timeout (seconds). |
maxRetries | int | No | Upstream retry count. |
ModelInfo
| Field | Type | Description |
|---|---|---|
maxTokens | int | Total context window. |
inputCostPerToken | float64 | Input token cost (dollars per token). |
outputCostPerToken | float64 | Output token cost (dollars per token). |
status
| Field | Type | Description |
|---|---|---|
syncedGateways[].gatewayName | string | Referenced gateway. |
syncedGateways[].synced | bool | Successful sync to LiteLLM. |
syncedGateways[].litellmModelRef | string | Name of the generated LiteLLMModel. |
syncedGateways[].lastSyncTime | string | RFC3339 timestamp. |
syncedGateways[].error | string | Most recent sync error, if any. |
conditions | []metav1.Condition | Standard conditions (Synced, Ready). |
Print columns
Model, Age.
Example
yaml
apiVersion: operator.palena.ai/v1alpha1
kind: PalenaModel
metadata:
name: gpt-4o-mini
namespace: palena
spec:
gatewayRefs:
- name: production
- name: staging
modelName: gpt-4o-mini
litellmParams:
model: openai/gpt-4o-mini
apiKeySecretRef:
name: openai-keys
key: apiKey
rpm: 10000
tpm: 2000000
modelInfo:
maxTokens: 128000
inputCostPerToken: 0.00000015
outputCostPerToken: 0.0000006
dataClassificationMaxLevel: 2Multi-gateway fan-out
When gatewayRefs lists more than one gateway, the controller creates one LiteLLMModel CR per gateway, each named <modelName>-<gatewayName>. Removing a gateway from the list cascades into deletion of the orphaned LiteLLMModel.