# Research: Which self-hosted AI gateway/proxy tool fits this effort's needs? **Question:** Which self-hosted AI gateway/proxy tool should front llama.cpp, given the requirements in issue #10 (Anthropic/OpenAI-compatible routing, per-workload virtual keys with separate usage views, a spend dashboard, custom cost-per-token pricing for a local model, docker-compose self-hosting alongside the existing stack, room to add backends later, and a plus for native queue/priority support relevant to #16)? **Answer: LiteLLM proxy.** It is the only candidate that meets every hard requirement out of the box, self-hosted, in its free/MIT tier. Portkey's open-source gateway is disqualified on the dashboard/virtual-key/budget requirement (those are cloud-only). Helicone is a weaker fit (maintenance mode, unclear virtual-key/custom-pricing story, feature-reduced self-host build, observability-first rather than budget/gateway-first). A hand-rolled nginx+script layer would mean re-building LiteLLM's virtual-key store, spend DB, dashboard, and Anthropic↔OpenAI translation from scratch — a maintenance trap, not a shortcut. ## Candidate: LiteLLM proxy **License / project health:** MIT-licensed, with a separate `enterprise/` subdirectory under its own license for a small set of add-on features (SSO, audit logs, guaranteed-capacity priority reservation — see below). Widely deployed, 100+ provider integrations. Source: [BerriAI/litellm LICENSE](https://raw.githubusercontent.com/BerriAI/litellm/main/LICENSE), [BerriAI/litellm GitHub repo](https://github.com/BerriAI/litellm). **Anthropic/OpenAI-compatible routing:** LiteLLM proxy exposes a unified `/v1/messages` endpoint that accepts Anthropic-format requests and translates them to whatever backend format the target model needs (and translates the response back), so Anthropic-format clients (coding CLIs) and OpenAI-format clients (Open WebUI) can both hit the same proxy against the same `model_list` entry pointing at llama.cpp's OpenAI-compatible `/v1/chat/completions`. This means llama.cpp's own native `/v1/messages` shim doesn't strictly need to be reached directly through the proxy — LiteLLM does its own Anthropic↔OpenAI translation in front of llama.cpp's OpenAI endpoint, which is one plausible wiring; routing straight through to llama.cpp's native shim as a passthrough is a second option worth checking at implementation time (issue #14/#15 territory, not this ticket). Source: [LiteLLM /v1/messages unified endpoint docs](https://docs.litellm.ai/docs/anthropic_unified/), [LiteLLM — Claude Code with non-Anthropic models](https://docs.litellm.ai/docs/tutorials/claude_non_anthropic_models). **Virtual keys / per-workload accounts:** `/key/generate` issues a virtual key with its own `max_budget`, `budget_duration`, and `tpm_limit`/`rpm_limit`. Keys can be owned by a `user_id` or a `team_id`, so each workload (Open WebUI, a coding CLI, Gitea code review, Paperless OCR, etc.) gets its own key with its own budget and its own spend record, queryable via `/key/info` and aggregated per team via `/team/info`. Spend is written to the `LiteLLM_VerificationTokenTable` and computed via LiteLLM's own `completion_cost()` on every call. Source: [LiteLLM — Virtual Keys](https://docs.litellm.ai/docs/proxy/virtual_keys). **Spend dashboard:** The Admin UI (`/ui`) ships in the open-source build — key/model management plus a Usage tab showing spend tracked per key/team. Enterprise adds SSO/SAML, audit logs, and a more advanced UI on top, but core spend-by-key visualization is not gated. Source: [LiteLLM — Proxy UI docs](https://docs.litellm.ai/docs/proxy/ui), cross-checked via [LiteLLM GitHub repo description](https://github.com/BerriAI/litellm) ("100+ LLM integrations, budgets, rate limits, logging, and virtual keys" in the free edition). **Custom cost-per-token pricing (shadow cloud-cost estimate):** Add a `model_info` block per model in `config.yaml`: ```yaml model_list: - model_name: my-local-model litellm_params: model: openai/local-model # or whatever provider shim fits llama.cpp api_base: http://llama-server:8080/v1 model_info: input_cost_per_token: 0.000001 output_cost_per_token: 0.000002 ``` This is exactly the mechanism #11 (reference cloud pricing) needs: once #11 picks a reference cloud model/price, its per-token rate goes straight into this block and LiteLLM computes "what this local usage would have cost" using its normal `completion_cost()` path — no separate cost-tracking code needed. Source: [LiteLLM — Custom pricing docs](https://docs.litellm.ai/docs/proxy/custom_pricing). **Docker-compose deployment:** The documented quickstart is a two-service compose (LiteLLM gateway + Postgres, Postgres storing keys/models/spend logs); Redis is optional and only needed for multi-instance state (rate limiting, cross-instance priority queueing) — a single-instance deployment alongside the existing llama.cpp/Open WebUI/Qdrant/Lazytainer stack doesn't require it. `LITELLM_SALT_KEY` must be set to something real before production use (it encrypts stored provider keys); the documented example otherwise uses `sk-1234` as a placeholder master key that must be replaced. A YAML-only, no-database mode exists but drops budget enforcement — not useful here since budgets/spend-per-key are a hard requirement. Source: [LiteLLM — Docker Quick Start](https://docs.litellm.ai/docs/proxy/docker_quick_start). **Room for more backends later:** LiteLLM's whole design point is a `model_list` of arbitrary provider entries behind one routing layer (100+ providers documented) — adding a second/third backend later is a config edit, not an architecture change. Source: [BerriAI/litellm GitHub repo](https://github.com/BerriAI/litellm). **Queuing/priority (relevant to #16, not decided here):** LiteLLM has an open-source (beta) request-prioritization scheduler: callers pass a `priority` field (lower number = higher priority) and a Router-level queue polls until a slot opens; multi-instance deployments need Redis to share queue state. This is a real, if beta-quality, building block for #16's priority queue and would mean #16 doesn't need a separate queuing component in front of the proxy. However: - The stricter **`priority_reservation`** feature (hard-reserving a % of TPM/RPM capacity per priority tier — not just soft-prioritizing) is gated behind the enterprise license. - The plain `priority` parameter had a real reported bug (leaking the `priority` field into the provider request, breaking calls to some providers) that maintainers closed as "not planned" rather than fixed — worth a smoke test against llama.cpp specifically before #16 relies on it, and worth treating the whole feature as "beta, verify before depending on it" rather than a settled capability. Sources: [LiteLLM — Request Prioritization (scheduler) docs](https://docs.litellm.ai/docs/scheduler), [BerriAI/litellm issue #7144 — "Priority feature is broken"](https://github.com/BerriAI/litellm/issues/7144) (closed not-planned), [BerriAI/litellm issue #6867 — scheduler polling bug](https://github.com/BerriAI/litellm/issues/6867), [BerriAI/litellm issue #13405 — feature request for priority-based request handling via API keys, i.e. the simpler ergonomic form isn't fully built yet either](https://github.com/BerriAI/litellm/issues/13405). ## Candidate: Portkey — disqualified Portkey's core gateway (`portkey-ai/gateway`) went fully open-source (Apache 2.0) and is self-hostable via Docker with routing, fallbacks, retries, load balancing, and a local logging console. But per the repo's own docs, **usage analytics/cost tracking, budget enforcement, and the spend dashboard are explicitly listed as "available in hosted and enterprise versions"** — i.e. they require Portkey's cloud control plane, not the self-hosted gateway alone. That fails this ticket's hard requirement for a self-hosted spend dashboard and per-key usage views, so Portkey is out regardless of its otherwise-solid routing feature set. Source: [Portkey-AI/gateway GitHub repo](https://github.com/portkey-ai/gateway). ## Candidate: Helicone — weaker fit Helicone is self-hostable via a documented docker-compose stack (web dashboard, a combined API+proxy service, Postgres, ClickHouse, MinIO for S3-compatible storage), Apache-2.0, so it can be run independent of Helicone's own cloud. But: - **Project status:** Mintlify acquired Helicone (2026-03-03) and the product is reported to be in maintenance mode; the repo is still active and Apache-2.0 so self-hosting isn't blocked, but it's a weaker bet for a component meant to grow (more backends, more workloads) over time. - **Feature parity gap:** the self-hosted docs explicitly note other providers (Vertex AI, Bedrock, Azure OpenAI) aren't supported in the self-hosted build the way they are in the cloud version — a signal the self-host path is the less-maintained one. - **Virtual keys / custom pricing:** not clearly documented for the self-hosted build in the pages checked — Helicone's primary framing is LLM *observability* (logging, tracing, cost dashboards computed from known-provider pricing tables) rather than a virtual-key issuing/budget gateway; injecting a custom per-token price for an unlisted local model isn't documented the way LiteLLM's `model_info.input_cost_per_token` is. - **Format translation:** self-host docs show separate `oai/` and `anthropic/` proxy paths rather than a documented single endpoint that translates Anthropic-format calls to an OpenAI-format backend the way LiteLLM's unified `/v1/messages` does. Sources: [Helicone — Docker Compose self-deploy docs](https://docs.helicone.ai/getting-started/self-deploy-docker), [Helicone — self-hosting launch announcement](https://www.helicone.ai/blog/self-hosting-launch), [Helicone/helicone GitHub repo](https://github.com/helicone/helicone). **Confidence note:** these self-host feature-gap claims come from docs-page summaries rather than a hands-on deployment; if Helicone is ever reconsidered, verify virtual-key/custom-pricing support directly against a running self-hosted instance rather than trusting this summary alone. ## Candidate: hand-rolled nginx + scripts — rejected as a maintenance trap A thin reverse proxy plus custom scripts could technically satisfy each bullet in isolation (issue an API key = generate a token and check it in an nginx `map`/Lua script; track spend = write to a DB on each request; dashboard = a small custom UI; custom pricing = a config file the script reads; Anthropic↔OpenAI translation = hand-written request/response transformers). But that's rebuilding LiteLLM's virtual-key store, spend ledger, dashboard, and format-translation layer from scratch, in a piece this repo would then own and maintain indefinitely, against a target (LiteLLM) that already does all of it, is MIT-licensed, and is a straightforward docker-compose service. Not adopted. ## Risks / gaps to carry into later tickets 1. **Lazytainer idle-suspend interaction (flagged in #9's "Not yet specified").** Lazytainer decides to stop `llama-server` based on network packet activity on its published port (`lazytainer.group.llamaserver.minPacketThreshold` / `inactiveTimeout` in this repo's `docker-compose.yml`). If LiteLLM proxy performs periodic background health checks against configured models (a common gateway behavior), that traffic could look like real usage to Lazytainer and prevent it from ever idling the container down. This needs to be checked against LiteLLM's actual health-check config (there are documented options to disable/tune background health checks) once the compose service is authored in #14, and verified on real hardware per #9's "Not yet specified" note on Lazytainer + multi-workload proxy interaction. 2. **llama.cpp's native `/v1/messages` shim vs. LiteLLM's own Anthropic translation.** Two wiring options exist (LiteLLM translates Anthropic→OpenAI itself and hits llama.cpp's OpenAI endpoint; or LiteLLM passes Anthropic-format requests straight through to llama.cpp's own native shim). This ticket confirms both are plausible per LiteLLM's docs but doesn't pick one — that's implementation detail for #14/#15, and should be smoke-tested against the actual coding-CLI flows in `docs/coding-cli-setup.md` once wired. 3. **Priority queueing for #16 is a real but beta/rough-edged LiteLLM feature**, with at least one reported-and-declined bug in the exact `priority` mechanism, and the stronger reserved-capacity variant is enterprise-gated. #16 should treat LiteLLM's scheduler as a starting point to validate hands-on, not an assumed solved problem — if it doesn't hold up under test, #16 may need a lightweight queuing shim in front of the proxy (e.g. a small request-queue sidecar) rather than reworking the whole gateway choice. 4. **Redis need is deferred, not eliminated.** A single-instance LiteLLM deployment (the right size for this effort) doesn't need Redis for virtual keys/spend/dashboard, but the priority scheduler's multi-instance behavior and rate-limit sharing do use it — if #16 ends up needing Redis-backed prioritization even in a single-instance deployment, add a `redis` service to the compose file at that point; no need to provision it speculatively now. ## Bottom line for the wayfinder map - Adopt **LiteLLM proxy** (MIT-licensed, `BerriAI/litellm`) as the gateway/proxy in front of llama.cpp for this effort. - It meets every hard requirement in #10 in its free/OSS build: virtual keys with per-key spend, a self-hosted Admin UI spend dashboard, config-driven custom per-token pricing (feeds #11 directly), a documented two-service docker-compose deployment, and an arbitrary-provider `model_list` that keeps future backends a config change away. - Its beta priority/queueing feature is a promising but unproven fit for #16 — validate it hands-on rather than assuming it's settled. - Portkey's self-hosted OSS gateway is disqualified (no self-hosted dashboard/budgets). Helicone is a workable but weaker fallback (maintenance-mode signal, feature-reduced self-host build, less clearly documented virtual-key/custom-pricing support) if LiteLLM turns out to be a poor fit during implementation.