From a1c37de6b4b8cdf9e9a5f8780319144e3ff79c97 Mon Sep 17 00:00:00 2001 From: Haylan Date: Tue, 25 Aug 2026 06:45:10 +0200 Subject: [PATCH] =?UTF-8?q?research:=20self-hosted=20AI=20gateway/proxy=20?= =?UTF-8?q?tool=20choice=20=E2=80=94=20recommend=20LiteLLM?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Sonnet 5 --- docs/research/proxy-tool-choice.md | 241 +++++++++++++++++++++++++++++ 1 file changed, 241 insertions(+) create mode 100644 docs/research/proxy-tool-choice.md diff --git a/docs/research/proxy-tool-choice.md b/docs/research/proxy-tool-choice.md new file mode 100644 index 0000000..b25897d --- /dev/null +++ b/docs/research/proxy-tool-choice.md @@ -0,0 +1,241 @@ +# Research: Which self-hosted AI gateway/proxy tool fits this effort's needs? + +**Question:** Which self-hosted AI gateway/proxy tool should front llama.cpp, +given the requirements in issue #10 (Anthropic/OpenAI-compatible routing, +per-workload virtual keys with separate usage views, a spend dashboard, +custom cost-per-token pricing for a local model, docker-compose +self-hosting alongside the existing stack, room to add backends later, and +a plus for native queue/priority support relevant to #16)? + +**Answer: LiteLLM proxy.** It is the only candidate that meets every +hard requirement out of the box, self-hosted, in its free/MIT tier. Portkey's +open-source gateway is disqualified on the dashboard/virtual-key/budget +requirement (those are cloud-only). Helicone is a weaker fit (maintenance +mode, unclear virtual-key/custom-pricing story, feature-reduced self-host +build, observability-first rather than budget/gateway-first). A hand-rolled +nginx+script layer would mean re-building LiteLLM's virtual-key store, spend +DB, dashboard, and Anthropic↔OpenAI translation from scratch — a maintenance +trap, not a shortcut. + +## Candidate: LiteLLM proxy + +**License / project health:** MIT-licensed, with a separate `enterprise/` +subdirectory under its own license for a small set of add-on features (SSO, +audit logs, guaranteed-capacity priority reservation — see below). Widely +deployed, 100+ provider integrations. +Source: [BerriAI/litellm LICENSE](https://raw.githubusercontent.com/BerriAI/litellm/main/LICENSE), +[BerriAI/litellm GitHub repo](https://github.com/BerriAI/litellm). + +**Anthropic/OpenAI-compatible routing:** LiteLLM proxy exposes a unified +`/v1/messages` endpoint that accepts Anthropic-format requests and +translates them to whatever backend format the target model needs (and +translates the response back), so Anthropic-format clients (coding CLIs) +and OpenAI-format clients (Open WebUI) can both hit the same proxy against +the same `model_list` entry pointing at llama.cpp's OpenAI-compatible +`/v1/chat/completions`. This means llama.cpp's own native `/v1/messages` +shim doesn't strictly need to be reached directly through the proxy — LiteLLM +does its own Anthropic↔OpenAI translation in front of llama.cpp's OpenAI +endpoint, which is one plausible wiring; routing straight through to +llama.cpp's native shim as a passthrough is a second option worth checking +at implementation time (issue #14/#15 territory, not this ticket). +Source: [LiteLLM /v1/messages unified endpoint docs](https://docs.litellm.ai/docs/anthropic_unified/), +[LiteLLM — Claude Code with non-Anthropic models](https://docs.litellm.ai/docs/tutorials/claude_non_anthropic_models). + +**Virtual keys / per-workload accounts:** `/key/generate` issues a virtual +key with its own `max_budget`, `budget_duration`, and `tpm_limit`/`rpm_limit`. +Keys can be owned by a `user_id` or a `team_id`, so each workload (Open +WebUI, a coding CLI, Gitea code review, Paperless OCR, etc.) gets its own +key with its own budget and its own spend record, queryable via +`/key/info` and aggregated per team via `/team/info`. Spend is written to +the `LiteLLM_VerificationTokenTable` and computed via LiteLLM's own +`completion_cost()` on every call. +Source: [LiteLLM — Virtual Keys](https://docs.litellm.ai/docs/proxy/virtual_keys). + +**Spend dashboard:** The Admin UI (`/ui`) ships in the open-source build — +key/model management plus a Usage tab showing spend tracked per key/team. +Enterprise adds SSO/SAML, audit logs, and a more advanced UI on top, but +core spend-by-key visualization is not gated. +Source: [LiteLLM — Proxy UI docs](https://docs.litellm.ai/docs/proxy/ui), +cross-checked via [LiteLLM GitHub repo description](https://github.com/BerriAI/litellm) +("100+ LLM integrations, budgets, rate limits, logging, and virtual keys" in +the free edition). + +**Custom cost-per-token pricing (shadow cloud-cost estimate):** Add a +`model_info` block per model in `config.yaml`: + +```yaml +model_list: + - model_name: my-local-model + litellm_params: + model: openai/local-model # or whatever provider shim fits llama.cpp + api_base: http://llama-server:8080/v1 + model_info: + input_cost_per_token: 0.000001 + output_cost_per_token: 0.000002 +``` + +This is exactly the mechanism #11 (reference cloud pricing) needs: once #11 +picks a reference cloud model/price, its per-token rate goes straight into +this block and LiteLLM computes "what this local usage would have cost" +using its normal `completion_cost()` path — no separate cost-tracking code +needed. +Source: [LiteLLM — Custom pricing docs](https://docs.litellm.ai/docs/proxy/custom_pricing). + +**Docker-compose deployment:** The documented quickstart is a two-service +compose (LiteLLM gateway + Postgres, Postgres storing keys/models/spend +logs); Redis is optional and only needed for multi-instance state (rate +limiting, cross-instance priority queueing) — a single-instance deployment +alongside the existing llama.cpp/Open WebUI/Qdrant/Lazytainer stack doesn't +require it. `LITELLM_SALT_KEY` must be set to something real before +production use (it encrypts stored provider keys); the documented example +otherwise uses `sk-1234` as a placeholder master key that must be replaced. +A YAML-only, no-database mode exists but drops budget enforcement — not +useful here since budgets/spend-per-key are a hard requirement. +Source: [LiteLLM — Docker Quick Start](https://docs.litellm.ai/docs/proxy/docker_quick_start). + +**Room for more backends later:** LiteLLM's whole design point is a +`model_list` of arbitrary provider entries behind one routing layer (100+ +providers documented) — adding a second/third backend later is a config +edit, not an architecture change. +Source: [BerriAI/litellm GitHub repo](https://github.com/BerriAI/litellm). + +**Queuing/priority (relevant to #16, not decided here):** LiteLLM has an +open-source (beta) request-prioritization scheduler: callers pass a +`priority` field (lower number = higher priority) and a Router-level queue +polls until a slot opens; multi-instance deployments need Redis to share +queue state. This is a real, if beta-quality, building block for #16's +priority queue and would mean #16 doesn't need a separate queuing +component in front of the proxy. However: +- The stricter **`priority_reservation`** feature (hard-reserving a % of + TPM/RPM capacity per priority tier — not just soft-prioritizing) is + gated behind the enterprise license. +- The plain `priority` parameter had a real reported bug (leaking the + `priority` field into the provider request, breaking calls to some + providers) that maintainers closed as "not planned" rather than fixed — + worth a smoke test against llama.cpp specifically before #16 relies on + it, and worth treating the whole feature as "beta, verify before + depending on it" rather than a settled capability. +Sources: [LiteLLM — Request Prioritization (scheduler) docs](https://docs.litellm.ai/docs/scheduler), +[BerriAI/litellm issue #7144 — "Priority feature is broken"](https://github.com/BerriAI/litellm/issues/7144) +(closed not-planned), +[BerriAI/litellm issue #6867 — scheduler polling bug](https://github.com/BerriAI/litellm/issues/6867), +[BerriAI/litellm issue #13405 — feature request for priority-based request +handling via API keys, i.e. the simpler ergonomic form isn't fully built +yet either](https://github.com/BerriAI/litellm/issues/13405). + +## Candidate: Portkey — disqualified + +Portkey's core gateway (`portkey-ai/gateway`) went fully open-source +(Apache 2.0) and is self-hostable via Docker with routing, fallbacks, +retries, load balancing, and a local logging console. But per the repo's +own docs, **usage analytics/cost tracking, budget enforcement, and the +spend dashboard are explicitly listed as "available in hosted and +enterprise versions"** — i.e. they require Portkey's cloud control plane, +not the self-hosted gateway alone. That fails this ticket's hard +requirement for a self-hosted spend dashboard and per-key usage views, so +Portkey is out regardless of its otherwise-solid routing feature set. +Source: [Portkey-AI/gateway GitHub repo](https://github.com/portkey-ai/gateway). + +## Candidate: Helicone — weaker fit + +Helicone is self-hostable via a documented docker-compose stack (web +dashboard, a combined API+proxy service, Postgres, ClickHouse, MinIO for +S3-compatible storage), Apache-2.0, so it can be run independent of +Helicone's own cloud. But: +- **Project status:** Mintlify acquired Helicone (2026-03-03) and the + product is reported to be in maintenance mode; the repo is still active + and Apache-2.0 so self-hosting isn't blocked, but it's a weaker bet for + a component meant to grow (more backends, more workloads) over time. +- **Feature parity gap:** the self-hosted docs explicitly note other + providers (Vertex AI, Bedrock, Azure OpenAI) aren't supported in the + self-hosted build the way they are in the cloud version — a signal the + self-host path is the less-maintained one. +- **Virtual keys / custom pricing:** not clearly documented for the + self-hosted build in the pages checked — Helicone's primary framing is + LLM *observability* (logging, tracing, cost dashboards computed from + known-provider pricing tables) rather than a virtual-key issuing/budget + gateway; injecting a custom per-token price for an unlisted local model + isn't documented the way LiteLLM's `model_info.input_cost_per_token` is. +- **Format translation:** self-host docs show separate `oai/` and + `anthropic/` proxy paths rather than a documented single endpoint that + translates Anthropic-format calls to an OpenAI-format backend the way + LiteLLM's unified `/v1/messages` does. + +Sources: [Helicone — Docker Compose self-deploy docs](https://docs.helicone.ai/getting-started/self-deploy-docker), +[Helicone — self-hosting launch announcement](https://www.helicone.ai/blog/self-hosting-launch), +[Helicone/helicone GitHub repo](https://github.com/helicone/helicone). +**Confidence note:** these self-host feature-gap claims come from +docs-page summaries rather than a hands-on deployment; if Helicone is +ever reconsidered, verify virtual-key/custom-pricing support directly +against a running self-hosted instance rather than trusting this summary +alone. + +## Candidate: hand-rolled nginx + scripts — rejected as a maintenance trap + +A thin reverse proxy plus custom scripts could technically satisfy each +bullet in isolation (issue an API key = generate a token and check it in +an nginx `map`/Lua script; track spend = write to a DB on each request; +dashboard = a small custom UI; custom pricing = a config file the script +reads; Anthropic↔OpenAI translation = hand-written request/response +transformers). But that's rebuilding LiteLLM's virtual-key store, spend +ledger, dashboard, and format-translation layer from scratch, in a +piece this repo would then own and maintain indefinitely, against a +target (LiteLLM) that already does all of it, is MIT-licensed, and is +a straightforward docker-compose service. Not adopted. + +## Risks / gaps to carry into later tickets + +1. **Lazytainer idle-suspend interaction (flagged in #9's "Not yet + specified").** Lazytainer decides to stop `llama-server` based on + network packet activity on its published port + (`lazytainer.group.llamaserver.minPacketThreshold` / + `inactiveTimeout` in this repo's `docker-compose.yml`). If LiteLLM + proxy performs periodic background health checks against configured + models (a common gateway behavior), that traffic could look like + real usage to Lazytainer and prevent it from ever idling the + container down. This needs to be checked against LiteLLM's actual + health-check config (there are documented options to disable/tune + background health checks) once the compose service is authored in + #14, and verified on real hardware per #9's "Not yet specified" note + on Lazytainer + multi-workload proxy interaction. +2. **llama.cpp's native `/v1/messages` shim vs. LiteLLM's own Anthropic + translation.** Two wiring options exist (LiteLLM translates + Anthropic→OpenAI itself and hits llama.cpp's OpenAI endpoint; or + LiteLLM passes Anthropic-format requests straight through to + llama.cpp's own native shim). This ticket confirms both are + plausible per LiteLLM's docs but doesn't pick one — that's + implementation detail for #14/#15, and should be smoke-tested against + the actual coding-CLI flows in `docs/coding-cli-setup.md` once wired. +3. **Priority queueing for #16 is a real but beta/rough-edged LiteLLM + feature**, with at least one reported-and-declined bug in the exact + `priority` mechanism, and the stronger reserved-capacity variant is + enterprise-gated. #16 should treat LiteLLM's scheduler as a + starting point to validate hands-on, not an assumed solved problem — + if it doesn't hold up under test, #16 may need a lightweight queuing + shim in front of the proxy (e.g. a small request-queue sidecar) rather + than reworking the whole gateway choice. +4. **Redis need is deferred, not eliminated.** A single-instance LiteLLM + deployment (the right size for this effort) doesn't need Redis for + virtual keys/spend/dashboard, but the priority scheduler's + multi-instance behavior and rate-limit sharing do use it — if #16 + ends up needing Redis-backed prioritization even in a single-instance + deployment, add a `redis` service to the compose file at that point; + no need to provision it speculatively now. + +## Bottom line for the wayfinder map + +- Adopt **LiteLLM proxy** (MIT-licensed, `BerriAI/litellm`) as the + gateway/proxy in front of llama.cpp for this effort. +- It meets every hard requirement in #10 in its free/OSS build: virtual + keys with per-key spend, a self-hosted Admin UI spend dashboard, + config-driven custom per-token pricing (feeds #11 directly), a + documented two-service docker-compose deployment, and an + arbitrary-provider `model_list` that keeps future backends a config + change away. +- Its beta priority/queueing feature is a promising but unproven fit for + #16 — validate it hands-on rather than assuming it's settled. +- Portkey's self-hosted OSS gateway is disqualified (no self-hosted + dashboard/budgets). Helicone is a workable but weaker fallback + (maintenance-mode signal, feature-reduced self-host build, less clearly + documented virtual-key/custom-pricing support) if LiteLLM turns out to + be a poor fit during implementation.