What is Palena?
Palena is a Kubernetes operator that composes CloudNativePG, LiteLLM, Langfuse, and a LibreChat-based UI into a single declarative AI platform. You write one high-level custom resource; Palena translates it into a fully wired production stack.
It's developed by bitkaio LLC and released under the Apache License, Version 2.0. Contributions are accepted under the same Apache 2.0 terms plus a signed Contributor License Agreement.
What problem does it solve?
Running an enterprise LLM platform on Kubernetes is a composition problem. You need:
- A gateway that handles auth, rate limiting, routing, and provider failover (LiteLLM).
- A database for the gateway's state, key store, and audit log (PostgreSQL, ideally HA).
- A cache for rate limits and model routing decisions (Redis).
- A chat UI for end users (LibreChat, Open WebUI, etc.).
- Observability — request-level tracing, cost tracking, evaluation (Langfuse).
- Tool-use infrastructure — MCP servers for websearch, retrieval, code execution, etc.
Each of these has its own Helm chart or operator, its own config file, its own secret convention. Wiring them together correctly — and keeping them wired across upgrades — is a full-time job.
Palena collapses that wiring into five Kubernetes CRDs. You describe what you want; the operator figures out how to get there.
What the operator creates
One PalenaGateway CR + supporting CRs results in:
| You write | Palena creates |
|---|---|
PalenaGateway | CNPG Cluster + Redis + LiteLLMInstance CR |
PalenaModel (≥1) | One LiteLLMModel CR per referenced gateway |
PalenaUI | LibreChat Deployment + MongoDB + MeiliSearch |
PalenaObservability | LangfuseInstance CR + callback wired into LiteLLM |
PalenaMCPServer | websearch MCP + SearXNG + Chromium + Presidio + FlashRank |
Every created resource gets an owner reference, so kubectl delete palenagateway production reliably tears down the entire stack.
Design principles
- One source of truth. Every knob lives in the CR spec — no side-channel ConfigMaps, no templated Helm values to diverge.
- BYO secrets. All credentials come from Secrets you control. Palena never stores plaintext anywhere.
- Everything owned. Operator sets owner references on every created resource for GC.
- Ingress + TLS built in. All public-facing components accept an
ingressfield with optional cert-manager wiring. - Observability-ready. Drop a
PalenaObservabilityCR and every gateway request gets traced in Langfuse. The callback is wired for you.
When not to use Palena
Palena is opinionated. It wires LiteLLM as the gateway, LibreChat as the UI, and Langfuse as the observability backend. If your architecture already standardizes on different components (for example, Open WebUI instead of LibreChat, or Helicone instead of Langfuse), you may find the compose layer fighting you.
Palena is also aimed at full-stack platform deployments. If you only need a LiteLLM gateway with no UI, no tracing, and no MCP servers, installing the upstream LiteLLM Operator directly is simpler.
What's next
- Architecture — the composition graph, reference graph, and deletion order
- Prerequisites — upstream operators you need before installing Palena
- Installation — OLM, Helm, or raw manifests
- Quick start — a working gateway + model + UI in under 10 minutes