14 KiB
Research: Which self-hosted AI gateway/proxy tool fits this effort's needs?
Question: Which self-hosted AI gateway/proxy tool should front llama.cpp, given the requirements in issue #10 (Anthropic/OpenAI-compatible routing, per-workload virtual keys with separate usage views, a spend dashboard, custom cost-per-token pricing for a local model, docker-compose self-hosting alongside the existing stack, room to add backends later, and a plus for native queue/priority support relevant to #16)?
Answer: LiteLLM proxy. It is the only candidate that meets every hard requirement out of the box, self-hosted, in its free/MIT tier. Portkey's open-source gateway is disqualified on the dashboard/virtual-key/budget requirement (those are cloud-only). Helicone is a weaker fit (maintenance mode, unclear virtual-key/custom-pricing story, feature-reduced self-host build, observability-first rather than budget/gateway-first). A hand-rolled nginx+script layer would mean re-building LiteLLM's virtual-key store, spend DB, dashboard, and Anthropic↔OpenAI translation from scratch — a maintenance trap, not a shortcut.
Candidate: LiteLLM proxy
License / project health: MIT-licensed, with a separate enterprise/
subdirectory under its own license for a small set of add-on features (SSO,
audit logs, guaranteed-capacity priority reservation — see below). Widely
deployed, 100+ provider integrations.
Source: BerriAI/litellm LICENSE,
BerriAI/litellm GitHub repo.
Anthropic/OpenAI-compatible routing: LiteLLM proxy exposes a unified
/v1/messages endpoint that accepts Anthropic-format requests and
translates them to whatever backend format the target model needs (and
translates the response back), so Anthropic-format clients (coding CLIs)
and OpenAI-format clients (Open WebUI) can both hit the same proxy against
the same model_list entry pointing at llama.cpp's OpenAI-compatible
/v1/chat/completions. This means llama.cpp's own native /v1/messages
shim doesn't strictly need to be reached directly through the proxy — LiteLLM
does its own Anthropic↔OpenAI translation in front of llama.cpp's OpenAI
endpoint, which is one plausible wiring; routing straight through to
llama.cpp's native shim as a passthrough is a second option worth checking
at implementation time (issue #14/#15 territory, not this ticket).
Source: LiteLLM /v1/messages unified endpoint docs,
LiteLLM — Claude Code with non-Anthropic models.
Virtual keys / per-workload accounts: /key/generate issues a virtual
key with its own max_budget, budget_duration, and tpm_limit/rpm_limit.
Keys can be owned by a user_id or a team_id, so each workload (Open
WebUI, a coding CLI, Gitea code review, Paperless OCR, etc.) gets its own
key with its own budget and its own spend record, queryable via
/key/info and aggregated per team via /team/info. Spend is written to
the LiteLLM_VerificationTokenTable and computed via LiteLLM's own
completion_cost() on every call.
Source: LiteLLM — Virtual Keys.
Spend dashboard: The Admin UI (/ui) ships in the open-source build —
key/model management plus a Usage tab showing spend tracked per key/team.
Enterprise adds SSO/SAML, audit logs, and a more advanced UI on top, but
core spend-by-key visualization is not gated.
Source: LiteLLM — Proxy UI docs,
cross-checked via LiteLLM GitHub repo description
("100+ LLM integrations, budgets, rate limits, logging, and virtual keys" in
the free edition).
Custom cost-per-token pricing (shadow cloud-cost estimate): Add a
model_info block per model in config.yaml:
model_list:
- model_name: my-local-model
litellm_params:
model: openai/local-model # or whatever provider shim fits llama.cpp
api_base: http://llama-server:8080/v1
model_info:
input_cost_per_token: 0.000001
output_cost_per_token: 0.000002
This is exactly the mechanism #11 (reference cloud pricing) needs: once #11
picks a reference cloud model/price, its per-token rate goes straight into
this block and LiteLLM computes "what this local usage would have cost"
using its normal completion_cost() path — no separate cost-tracking code
needed.
Source: LiteLLM — Custom pricing docs.
Docker-compose deployment: The documented quickstart is a two-service
compose (LiteLLM gateway + Postgres, Postgres storing keys/models/spend
logs); Redis is optional and only needed for multi-instance state (rate
limiting, cross-instance priority queueing) — a single-instance deployment
alongside the existing llama.cpp/Open WebUI/Qdrant/Lazytainer stack doesn't
require it. LITELLM_SALT_KEY must be set to something real before
production use (it encrypts stored provider keys); the documented example
otherwise uses sk-1234 as a placeholder master key that must be replaced.
A YAML-only, no-database mode exists but drops budget enforcement — not
useful here since budgets/spend-per-key are a hard requirement.
Source: LiteLLM — Docker Quick Start.
Room for more backends later: LiteLLM's whole design point is a
model_list of arbitrary provider entries behind one routing layer (100+
providers documented) — adding a second/third backend later is a config
edit, not an architecture change.
Source: BerriAI/litellm GitHub repo.
Queuing/priority (relevant to #16, not decided here): LiteLLM has an
open-source (beta) request-prioritization scheduler: callers pass a
priority field (lower number = higher priority) and a Router-level queue
polls until a slot opens; multi-instance deployments need Redis to share
queue state. This is a real, if beta-quality, building block for #16's
priority queue and would mean #16 doesn't need a separate queuing
component in front of the proxy. However:
- The stricter
priority_reservationfeature (hard-reserving a % of TPM/RPM capacity per priority tier — not just soft-prioritizing) is gated behind the enterprise license. - The plain
priorityparameter had a real reported bug (leaking thepriorityfield into the provider request, breaking calls to some providers) that maintainers closed as "not planned" rather than fixed — worth a smoke test against llama.cpp specifically before #16 relies on it, and worth treating the whole feature as "beta, verify before depending on it" rather than a settled capability. Sources: LiteLLM — Request Prioritization (scheduler) docs, BerriAI/litellm issue #7144 — "Priority feature is broken" (closed not-planned), BerriAI/litellm issue #6867 — scheduler polling bug, BerriAI/litellm issue #13405 — feature request for priority-based request handling via API keys, i.e. the simpler ergonomic form isn't fully built yet either.
Candidate: Portkey — disqualified
Portkey's core gateway (portkey-ai/gateway) went fully open-source
(Apache 2.0) and is self-hostable via Docker with routing, fallbacks,
retries, load balancing, and a local logging console. But per the repo's
own docs, usage analytics/cost tracking, budget enforcement, and the
spend dashboard are explicitly listed as "available in hosted and
enterprise versions" — i.e. they require Portkey's cloud control plane,
not the self-hosted gateway alone. That fails this ticket's hard
requirement for a self-hosted spend dashboard and per-key usage views, so
Portkey is out regardless of its otherwise-solid routing feature set.
Source: Portkey-AI/gateway GitHub repo.
Candidate: Helicone — weaker fit
Helicone is self-hostable via a documented docker-compose stack (web dashboard, a combined API+proxy service, Postgres, ClickHouse, MinIO for S3-compatible storage), Apache-2.0, so it can be run independent of Helicone's own cloud. But:
- Project status: Mintlify acquired Helicone (2026-03-03) and the product is reported to be in maintenance mode; the repo is still active and Apache-2.0 so self-hosting isn't blocked, but it's a weaker bet for a component meant to grow (more backends, more workloads) over time.
- Feature parity gap: the self-hosted docs explicitly note other providers (Vertex AI, Bedrock, Azure OpenAI) aren't supported in the self-hosted build the way they are in the cloud version — a signal the self-host path is the less-maintained one.
- Virtual keys / custom pricing: not clearly documented for the
self-hosted build in the pages checked — Helicone's primary framing is
LLM observability (logging, tracing, cost dashboards computed from
known-provider pricing tables) rather than a virtual-key issuing/budget
gateway; injecting a custom per-token price for an unlisted local model
isn't documented the way LiteLLM's
model_info.input_cost_per_tokenis. - Format translation: self-host docs show separate
oai/andanthropic/proxy paths rather than a documented single endpoint that translates Anthropic-format calls to an OpenAI-format backend the way LiteLLM's unified/v1/messagesdoes.
Sources: Helicone — Docker Compose self-deploy docs, Helicone — self-hosting launch announcement, Helicone/helicone GitHub repo. Confidence note: these self-host feature-gap claims come from docs-page summaries rather than a hands-on deployment; if Helicone is ever reconsidered, verify virtual-key/custom-pricing support directly against a running self-hosted instance rather than trusting this summary alone.
Candidate: hand-rolled nginx + scripts — rejected as a maintenance trap
A thin reverse proxy plus custom scripts could technically satisfy each
bullet in isolation (issue an API key = generate a token and check it in
an nginx map/Lua script; track spend = write to a DB on each request;
dashboard = a small custom UI; custom pricing = a config file the script
reads; Anthropic↔OpenAI translation = hand-written request/response
transformers). But that's rebuilding LiteLLM's virtual-key store, spend
ledger, dashboard, and format-translation layer from scratch, in a
piece this repo would then own and maintain indefinitely, against a
target (LiteLLM) that already does all of it, is MIT-licensed, and is
a straightforward docker-compose service. Not adopted.
Risks / gaps to carry into later tickets
- Lazytainer idle-suspend interaction (flagged in #9's "Not yet
specified"). Lazytainer decides to stop
llama-serverbased on network packet activity on its published port (lazytainer.group.llamaserver.minPacketThreshold/inactiveTimeoutin this repo'sdocker-compose.yml). If LiteLLM proxy performs periodic background health checks against configured models (a common gateway behavior), that traffic could look like real usage to Lazytainer and prevent it from ever idling the container down. This needs to be checked against LiteLLM's actual health-check config (there are documented options to disable/tune background health checks) once the compose service is authored in #14, and verified on real hardware per #9's "Not yet specified" note on Lazytainer + multi-workload proxy interaction. - llama.cpp's native
/v1/messagesshim vs. LiteLLM's own Anthropic translation. Two wiring options exist (LiteLLM translates Anthropic→OpenAI itself and hits llama.cpp's OpenAI endpoint; or LiteLLM passes Anthropic-format requests straight through to llama.cpp's own native shim). This ticket confirms both are plausible per LiteLLM's docs but doesn't pick one — that's implementation detail for #14/#15, and should be smoke-tested against the actual coding-CLI flows indocs/coding-cli-setup.mdonce wired. - Priority queueing for #16 is a real but beta/rough-edged LiteLLM
feature, with at least one reported-and-declined bug in the exact
prioritymechanism, and the stronger reserved-capacity variant is enterprise-gated. #16 should treat LiteLLM's scheduler as a starting point to validate hands-on, not an assumed solved problem — if it doesn't hold up under test, #16 may need a lightweight queuing shim in front of the proxy (e.g. a small request-queue sidecar) rather than reworking the whole gateway choice. - Redis need is deferred, not eliminated. A single-instance LiteLLM
deployment (the right size for this effort) doesn't need Redis for
virtual keys/spend/dashboard, but the priority scheduler's
multi-instance behavior and rate-limit sharing do use it — if #16
ends up needing Redis-backed prioritization even in a single-instance
deployment, add a
redisservice to the compose file at that point; no need to provision it speculatively now.
Bottom line for the wayfinder map
- Adopt LiteLLM proxy (MIT-licensed,
BerriAI/litellm) as the gateway/proxy in front of llama.cpp for this effort. - It meets every hard requirement in #10 in its free/OSS build: virtual
keys with per-key spend, a self-hosted Admin UI spend dashboard,
config-driven custom per-token pricing (feeds #11 directly), a
documented two-service docker-compose deployment, and an
arbitrary-provider
model_listthat keeps future backends a config change away. - Its beta priority/queueing feature is a promising but unproven fit for #16 — validate it hands-on rather than assuming it's settled.
- Portkey's self-hosted OSS gateway is disqualified (no self-hosted dashboard/budgets). Helicone is a workable but weaker fallback (maintenance-mode signal, feature-reduced self-host build, less clearly documented virtual-key/custom-pricing support) if LiteLLM turns out to be a poor fit during implementation.