Files
LLM-Server/docs/research/proxy-tool-choice.md
T

14 KiB

Research: Which self-hosted AI gateway/proxy tool fits this effort's needs?

Question: Which self-hosted AI gateway/proxy tool should front llama.cpp, given the requirements in issue #10 (Anthropic/OpenAI-compatible routing, per-workload virtual keys with separate usage views, a spend dashboard, custom cost-per-token pricing for a local model, docker-compose self-hosting alongside the existing stack, room to add backends later, and a plus for native queue/priority support relevant to #16)?

Answer: LiteLLM proxy. It is the only candidate that meets every hard requirement out of the box, self-hosted, in its free/MIT tier. Portkey's open-source gateway is disqualified on the dashboard/virtual-key/budget requirement (those are cloud-only). Helicone is a weaker fit (maintenance mode, unclear virtual-key/custom-pricing story, feature-reduced self-host build, observability-first rather than budget/gateway-first). A hand-rolled nginx+script layer would mean re-building LiteLLM's virtual-key store, spend DB, dashboard, and Anthropic↔OpenAI translation from scratch — a maintenance trap, not a shortcut.

Candidate: LiteLLM proxy

License / project health: MIT-licensed, with a separate enterprise/ subdirectory under its own license for a small set of add-on features (SSO, audit logs, guaranteed-capacity priority reservation — see below). Widely deployed, 100+ provider integrations. Source: BerriAI/litellm LICENSE, BerriAI/litellm GitHub repo.

Anthropic/OpenAI-compatible routing: LiteLLM proxy exposes a unified /v1/messages endpoint that accepts Anthropic-format requests and translates them to whatever backend format the target model needs (and translates the response back), so Anthropic-format clients (coding CLIs) and OpenAI-format clients (Open WebUI) can both hit the same proxy against the same model_list entry pointing at llama.cpp's OpenAI-compatible /v1/chat/completions. This means llama.cpp's own native /v1/messages shim doesn't strictly need to be reached directly through the proxy — LiteLLM does its own Anthropic↔OpenAI translation in front of llama.cpp's OpenAI endpoint, which is one plausible wiring; routing straight through to llama.cpp's native shim as a passthrough is a second option worth checking at implementation time (issue #14/#15 territory, not this ticket). Source: LiteLLM /v1/messages unified endpoint docs, LiteLLM — Claude Code with non-Anthropic models.

Virtual keys / per-workload accounts: /key/generate issues a virtual key with its own max_budget, budget_duration, and tpm_limit/rpm_limit. Keys can be owned by a user_id or a team_id, so each workload (Open WebUI, a coding CLI, Gitea code review, Paperless OCR, etc.) gets its own key with its own budget and its own spend record, queryable via /key/info and aggregated per team via /team/info. Spend is written to the LiteLLM_VerificationTokenTable and computed via LiteLLM's own completion_cost() on every call. Source: LiteLLM — Virtual Keys.

Spend dashboard: The Admin UI (/ui) ships in the open-source build — key/model management plus a Usage tab showing spend tracked per key/team. Enterprise adds SSO/SAML, audit logs, and a more advanced UI on top, but core spend-by-key visualization is not gated. Source: LiteLLM — Proxy UI docs, cross-checked via LiteLLM GitHub repo description ("100+ LLM integrations, budgets, rate limits, logging, and virtual keys" in the free edition).

Custom cost-per-token pricing (shadow cloud-cost estimate): Add a model_info block per model in config.yaml:

model_list:
  - model_name: my-local-model
    litellm_params:
      model: openai/local-model   # or whatever provider shim fits llama.cpp
      api_base: http://llama-server:8080/v1
    model_info:
      input_cost_per_token: 0.000001
      output_cost_per_token: 0.000002

This is exactly the mechanism #11 (reference cloud pricing) needs: once #11 picks a reference cloud model/price, its per-token rate goes straight into this block and LiteLLM computes "what this local usage would have cost" using its normal completion_cost() path — no separate cost-tracking code needed. Source: LiteLLM — Custom pricing docs.

Docker-compose deployment: The documented quickstart is a two-service compose (LiteLLM gateway + Postgres, Postgres storing keys/models/spend logs); Redis is optional and only needed for multi-instance state (rate limiting, cross-instance priority queueing) — a single-instance deployment alongside the existing llama.cpp/Open WebUI/Qdrant/Lazytainer stack doesn't require it. LITELLM_SALT_KEY must be set to something real before production use (it encrypts stored provider keys); the documented example otherwise uses sk-1234 as a placeholder master key that must be replaced. A YAML-only, no-database mode exists but drops budget enforcement — not useful here since budgets/spend-per-key are a hard requirement. Source: LiteLLM — Docker Quick Start.

Room for more backends later: LiteLLM's whole design point is a model_list of arbitrary provider entries behind one routing layer (100+ providers documented) — adding a second/third backend later is a config edit, not an architecture change. Source: BerriAI/litellm GitHub repo.

Queuing/priority (relevant to #16, not decided here): LiteLLM has an open-source (beta) request-prioritization scheduler: callers pass a priority field (lower number = higher priority) and a Router-level queue polls until a slot opens; multi-instance deployments need Redis to share queue state. This is a real, if beta-quality, building block for #16's priority queue and would mean #16 doesn't need a separate queuing component in front of the proxy. However:

Candidate: Portkey — disqualified

Portkey's core gateway (portkey-ai/gateway) went fully open-source (Apache 2.0) and is self-hostable via Docker with routing, fallbacks, retries, load balancing, and a local logging console. But per the repo's own docs, usage analytics/cost tracking, budget enforcement, and the spend dashboard are explicitly listed as "available in hosted and enterprise versions" — i.e. they require Portkey's cloud control plane, not the self-hosted gateway alone. That fails this ticket's hard requirement for a self-hosted spend dashboard and per-key usage views, so Portkey is out regardless of its otherwise-solid routing feature set. Source: Portkey-AI/gateway GitHub repo.

Candidate: Helicone — weaker fit

Helicone is self-hostable via a documented docker-compose stack (web dashboard, a combined API+proxy service, Postgres, ClickHouse, MinIO for S3-compatible storage), Apache-2.0, so it can be run independent of Helicone's own cloud. But:

  • Project status: Mintlify acquired Helicone (2026-03-03) and the product is reported to be in maintenance mode; the repo is still active and Apache-2.0 so self-hosting isn't blocked, but it's a weaker bet for a component meant to grow (more backends, more workloads) over time.
  • Feature parity gap: the self-hosted docs explicitly note other providers (Vertex AI, Bedrock, Azure OpenAI) aren't supported in the self-hosted build the way they are in the cloud version — a signal the self-host path is the less-maintained one.
  • Virtual keys / custom pricing: not clearly documented for the self-hosted build in the pages checked — Helicone's primary framing is LLM observability (logging, tracing, cost dashboards computed from known-provider pricing tables) rather than a virtual-key issuing/budget gateway; injecting a custom per-token price for an unlisted local model isn't documented the way LiteLLM's model_info.input_cost_per_token is.
  • Format translation: self-host docs show separate oai/ and anthropic/ proxy paths rather than a documented single endpoint that translates Anthropic-format calls to an OpenAI-format backend the way LiteLLM's unified /v1/messages does.

Sources: Helicone — Docker Compose self-deploy docs, Helicone — self-hosting launch announcement, Helicone/helicone GitHub repo. Confidence note: these self-host feature-gap claims come from docs-page summaries rather than a hands-on deployment; if Helicone is ever reconsidered, verify virtual-key/custom-pricing support directly against a running self-hosted instance rather than trusting this summary alone.

Candidate: hand-rolled nginx + scripts — rejected as a maintenance trap

A thin reverse proxy plus custom scripts could technically satisfy each bullet in isolation (issue an API key = generate a token and check it in an nginx map/Lua script; track spend = write to a DB on each request; dashboard = a small custom UI; custom pricing = a config file the script reads; Anthropic↔OpenAI translation = hand-written request/response transformers). But that's rebuilding LiteLLM's virtual-key store, spend ledger, dashboard, and format-translation layer from scratch, in a piece this repo would then own and maintain indefinitely, against a target (LiteLLM) that already does all of it, is MIT-licensed, and is a straightforward docker-compose service. Not adopted.

Risks / gaps to carry into later tickets

  1. Lazytainer idle-suspend interaction (flagged in #9's "Not yet specified"). Lazytainer decides to stop llama-server based on network packet activity on its published port (lazytainer.group.llamaserver.minPacketThreshold / inactiveTimeout in this repo's docker-compose.yml). If LiteLLM proxy performs periodic background health checks against configured models (a common gateway behavior), that traffic could look like real usage to Lazytainer and prevent it from ever idling the container down. This needs to be checked against LiteLLM's actual health-check config (there are documented options to disable/tune background health checks) once the compose service is authored in #14, and verified on real hardware per #9's "Not yet specified" note on Lazytainer + multi-workload proxy interaction.
  2. llama.cpp's native /v1/messages shim vs. LiteLLM's own Anthropic translation. Two wiring options exist (LiteLLM translates Anthropic→OpenAI itself and hits llama.cpp's OpenAI endpoint; or LiteLLM passes Anthropic-format requests straight through to llama.cpp's own native shim). This ticket confirms both are plausible per LiteLLM's docs but doesn't pick one — that's implementation detail for #14/#15, and should be smoke-tested against the actual coding-CLI flows in docs/coding-cli-setup.md once wired.
  3. Priority queueing for #16 is a real but beta/rough-edged LiteLLM feature, with at least one reported-and-declined bug in the exact priority mechanism, and the stronger reserved-capacity variant is enterprise-gated. #16 should treat LiteLLM's scheduler as a starting point to validate hands-on, not an assumed solved problem — if it doesn't hold up under test, #16 may need a lightweight queuing shim in front of the proxy (e.g. a small request-queue sidecar) rather than reworking the whole gateway choice.
  4. Redis need is deferred, not eliminated. A single-instance LiteLLM deployment (the right size for this effort) doesn't need Redis for virtual keys/spend/dashboard, but the priority scheduler's multi-instance behavior and rate-limit sharing do use it — if #16 ends up needing Redis-backed prioritization even in a single-instance deployment, add a redis service to the compose file at that point; no need to provision it speculatively now.

Bottom line for the wayfinder map

  • Adopt LiteLLM proxy (MIT-licensed, BerriAI/litellm) as the gateway/proxy in front of llama.cpp for this effort.
  • It meets every hard requirement in #10 in its free/OSS build: virtual keys with per-key spend, a self-hosted Admin UI spend dashboard, config-driven custom per-token pricing (feeds #11 directly), a documented two-service docker-compose deployment, and an arbitrary-provider model_list that keeps future backends a config change away.
  • Its beta priority/queueing feature is a promising but unproven fit for #16 — validate it hands-on rather than assuming it's settled.
  • Portkey's self-hosted OSS gateway is disqualified (no self-hosted dashboard/budgets). Helicone is a workable but weaker fallback (maintenance-mode signal, feature-reduced self-host build, less clearly documented virtual-key/custom-pricing support) if LiteLLM turns out to be a poor fit during implementation.