242 lines
14 KiB
Markdown
242 lines
14 KiB
Markdown
# Research: Which self-hosted AI gateway/proxy tool fits this effort's needs?
|
|
|
|
**Question:** Which self-hosted AI gateway/proxy tool should front llama.cpp,
|
|
given the requirements in issue #10 (Anthropic/OpenAI-compatible routing,
|
|
per-workload virtual keys with separate usage views, a spend dashboard,
|
|
custom cost-per-token pricing for a local model, docker-compose
|
|
self-hosting alongside the existing stack, room to add backends later, and
|
|
a plus for native queue/priority support relevant to #16)?
|
|
|
|
**Answer: LiteLLM proxy.** It is the only candidate that meets every
|
|
hard requirement out of the box, self-hosted, in its free/MIT tier. Portkey's
|
|
open-source gateway is disqualified on the dashboard/virtual-key/budget
|
|
requirement (those are cloud-only). Helicone is a weaker fit (maintenance
|
|
mode, unclear virtual-key/custom-pricing story, feature-reduced self-host
|
|
build, observability-first rather than budget/gateway-first). A hand-rolled
|
|
nginx+script layer would mean re-building LiteLLM's virtual-key store, spend
|
|
DB, dashboard, and Anthropic↔OpenAI translation from scratch — a maintenance
|
|
trap, not a shortcut.
|
|
|
|
## Candidate: LiteLLM proxy
|
|
|
|
**License / project health:** MIT-licensed, with a separate `enterprise/`
|
|
subdirectory under its own license for a small set of add-on features (SSO,
|
|
audit logs, guaranteed-capacity priority reservation — see below). Widely
|
|
deployed, 100+ provider integrations.
|
|
Source: [BerriAI/litellm LICENSE](https://raw.githubusercontent.com/BerriAI/litellm/main/LICENSE),
|
|
[BerriAI/litellm GitHub repo](https://github.com/BerriAI/litellm).
|
|
|
|
**Anthropic/OpenAI-compatible routing:** LiteLLM proxy exposes a unified
|
|
`/v1/messages` endpoint that accepts Anthropic-format requests and
|
|
translates them to whatever backend format the target model needs (and
|
|
translates the response back), so Anthropic-format clients (coding CLIs)
|
|
and OpenAI-format clients (Open WebUI) can both hit the same proxy against
|
|
the same `model_list` entry pointing at llama.cpp's OpenAI-compatible
|
|
`/v1/chat/completions`. This means llama.cpp's own native `/v1/messages`
|
|
shim doesn't strictly need to be reached directly through the proxy — LiteLLM
|
|
does its own Anthropic↔OpenAI translation in front of llama.cpp's OpenAI
|
|
endpoint, which is one plausible wiring; routing straight through to
|
|
llama.cpp's native shim as a passthrough is a second option worth checking
|
|
at implementation time (issue #14/#15 territory, not this ticket).
|
|
Source: [LiteLLM /v1/messages unified endpoint docs](https://docs.litellm.ai/docs/anthropic_unified/),
|
|
[LiteLLM — Claude Code with non-Anthropic models](https://docs.litellm.ai/docs/tutorials/claude_non_anthropic_models).
|
|
|
|
**Virtual keys / per-workload accounts:** `/key/generate` issues a virtual
|
|
key with its own `max_budget`, `budget_duration`, and `tpm_limit`/`rpm_limit`.
|
|
Keys can be owned by a `user_id` or a `team_id`, so each workload (Open
|
|
WebUI, a coding CLI, Gitea code review, Paperless OCR, etc.) gets its own
|
|
key with its own budget and its own spend record, queryable via
|
|
`/key/info` and aggregated per team via `/team/info`. Spend is written to
|
|
the `LiteLLM_VerificationTokenTable` and computed via LiteLLM's own
|
|
`completion_cost()` on every call.
|
|
Source: [LiteLLM — Virtual Keys](https://docs.litellm.ai/docs/proxy/virtual_keys).
|
|
|
|
**Spend dashboard:** The Admin UI (`/ui`) ships in the open-source build —
|
|
key/model management plus a Usage tab showing spend tracked per key/team.
|
|
Enterprise adds SSO/SAML, audit logs, and a more advanced UI on top, but
|
|
core spend-by-key visualization is not gated.
|
|
Source: [LiteLLM — Proxy UI docs](https://docs.litellm.ai/docs/proxy/ui),
|
|
cross-checked via [LiteLLM GitHub repo description](https://github.com/BerriAI/litellm)
|
|
("100+ LLM integrations, budgets, rate limits, logging, and virtual keys" in
|
|
the free edition).
|
|
|
|
**Custom cost-per-token pricing (shadow cloud-cost estimate):** Add a
|
|
`model_info` block per model in `config.yaml`:
|
|
|
|
```yaml
|
|
model_list:
|
|
- model_name: my-local-model
|
|
litellm_params:
|
|
model: openai/local-model # or whatever provider shim fits llama.cpp
|
|
api_base: http://llama-server:8080/v1
|
|
model_info:
|
|
input_cost_per_token: 0.000001
|
|
output_cost_per_token: 0.000002
|
|
```
|
|
|
|
This is exactly the mechanism #11 (reference cloud pricing) needs: once #11
|
|
picks a reference cloud model/price, its per-token rate goes straight into
|
|
this block and LiteLLM computes "what this local usage would have cost"
|
|
using its normal `completion_cost()` path — no separate cost-tracking code
|
|
needed.
|
|
Source: [LiteLLM — Custom pricing docs](https://docs.litellm.ai/docs/proxy/custom_pricing).
|
|
|
|
**Docker-compose deployment:** The documented quickstart is a two-service
|
|
compose (LiteLLM gateway + Postgres, Postgres storing keys/models/spend
|
|
logs); Redis is optional and only needed for multi-instance state (rate
|
|
limiting, cross-instance priority queueing) — a single-instance deployment
|
|
alongside the existing llama.cpp/Open WebUI/Qdrant/Lazytainer stack doesn't
|
|
require it. `LITELLM_SALT_KEY` must be set to something real before
|
|
production use (it encrypts stored provider keys); the documented example
|
|
otherwise uses `sk-1234` as a placeholder master key that must be replaced.
|
|
A YAML-only, no-database mode exists but drops budget enforcement — not
|
|
useful here since budgets/spend-per-key are a hard requirement.
|
|
Source: [LiteLLM — Docker Quick Start](https://docs.litellm.ai/docs/proxy/docker_quick_start).
|
|
|
|
**Room for more backends later:** LiteLLM's whole design point is a
|
|
`model_list` of arbitrary provider entries behind one routing layer (100+
|
|
providers documented) — adding a second/third backend later is a config
|
|
edit, not an architecture change.
|
|
Source: [BerriAI/litellm GitHub repo](https://github.com/BerriAI/litellm).
|
|
|
|
**Queuing/priority (relevant to #16, not decided here):** LiteLLM has an
|
|
open-source (beta) request-prioritization scheduler: callers pass a
|
|
`priority` field (lower number = higher priority) and a Router-level queue
|
|
polls until a slot opens; multi-instance deployments need Redis to share
|
|
queue state. This is a real, if beta-quality, building block for #16's
|
|
priority queue and would mean #16 doesn't need a separate queuing
|
|
component in front of the proxy. However:
|
|
- The stricter **`priority_reservation`** feature (hard-reserving a % of
|
|
TPM/RPM capacity per priority tier — not just soft-prioritizing) is
|
|
gated behind the enterprise license.
|
|
- The plain `priority` parameter had a real reported bug (leaking the
|
|
`priority` field into the provider request, breaking calls to some
|
|
providers) that maintainers closed as "not planned" rather than fixed —
|
|
worth a smoke test against llama.cpp specifically before #16 relies on
|
|
it, and worth treating the whole feature as "beta, verify before
|
|
depending on it" rather than a settled capability.
|
|
Sources: [LiteLLM — Request Prioritization (scheduler) docs](https://docs.litellm.ai/docs/scheduler),
|
|
[BerriAI/litellm issue #7144 — "Priority feature is broken"](https://github.com/BerriAI/litellm/issues/7144)
|
|
(closed not-planned),
|
|
[BerriAI/litellm issue #6867 — scheduler polling bug](https://github.com/BerriAI/litellm/issues/6867),
|
|
[BerriAI/litellm issue #13405 — feature request for priority-based request
|
|
handling via API keys, i.e. the simpler ergonomic form isn't fully built
|
|
yet either](https://github.com/BerriAI/litellm/issues/13405).
|
|
|
|
## Candidate: Portkey — disqualified
|
|
|
|
Portkey's core gateway (`portkey-ai/gateway`) went fully open-source
|
|
(Apache 2.0) and is self-hostable via Docker with routing, fallbacks,
|
|
retries, load balancing, and a local logging console. But per the repo's
|
|
own docs, **usage analytics/cost tracking, budget enforcement, and the
|
|
spend dashboard are explicitly listed as "available in hosted and
|
|
enterprise versions"** — i.e. they require Portkey's cloud control plane,
|
|
not the self-hosted gateway alone. That fails this ticket's hard
|
|
requirement for a self-hosted spend dashboard and per-key usage views, so
|
|
Portkey is out regardless of its otherwise-solid routing feature set.
|
|
Source: [Portkey-AI/gateway GitHub repo](https://github.com/portkey-ai/gateway).
|
|
|
|
## Candidate: Helicone — weaker fit
|
|
|
|
Helicone is self-hostable via a documented docker-compose stack (web
|
|
dashboard, a combined API+proxy service, Postgres, ClickHouse, MinIO for
|
|
S3-compatible storage), Apache-2.0, so it can be run independent of
|
|
Helicone's own cloud. But:
|
|
- **Project status:** Mintlify acquired Helicone (2026-03-03) and the
|
|
product is reported to be in maintenance mode; the repo is still active
|
|
and Apache-2.0 so self-hosting isn't blocked, but it's a weaker bet for
|
|
a component meant to grow (more backends, more workloads) over time.
|
|
- **Feature parity gap:** the self-hosted docs explicitly note other
|
|
providers (Vertex AI, Bedrock, Azure OpenAI) aren't supported in the
|
|
self-hosted build the way they are in the cloud version — a signal the
|
|
self-host path is the less-maintained one.
|
|
- **Virtual keys / custom pricing:** not clearly documented for the
|
|
self-hosted build in the pages checked — Helicone's primary framing is
|
|
LLM *observability* (logging, tracing, cost dashboards computed from
|
|
known-provider pricing tables) rather than a virtual-key issuing/budget
|
|
gateway; injecting a custom per-token price for an unlisted local model
|
|
isn't documented the way LiteLLM's `model_info.input_cost_per_token` is.
|
|
- **Format translation:** self-host docs show separate `oai/` and
|
|
`anthropic/` proxy paths rather than a documented single endpoint that
|
|
translates Anthropic-format calls to an OpenAI-format backend the way
|
|
LiteLLM's unified `/v1/messages` does.
|
|
|
|
Sources: [Helicone — Docker Compose self-deploy docs](https://docs.helicone.ai/getting-started/self-deploy-docker),
|
|
[Helicone — self-hosting launch announcement](https://www.helicone.ai/blog/self-hosting-launch),
|
|
[Helicone/helicone GitHub repo](https://github.com/helicone/helicone).
|
|
**Confidence note:** these self-host feature-gap claims come from
|
|
docs-page summaries rather than a hands-on deployment; if Helicone is
|
|
ever reconsidered, verify virtual-key/custom-pricing support directly
|
|
against a running self-hosted instance rather than trusting this summary
|
|
alone.
|
|
|
|
## Candidate: hand-rolled nginx + scripts — rejected as a maintenance trap
|
|
|
|
A thin reverse proxy plus custom scripts could technically satisfy each
|
|
bullet in isolation (issue an API key = generate a token and check it in
|
|
an nginx `map`/Lua script; track spend = write to a DB on each request;
|
|
dashboard = a small custom UI; custom pricing = a config file the script
|
|
reads; Anthropic↔OpenAI translation = hand-written request/response
|
|
transformers). But that's rebuilding LiteLLM's virtual-key store, spend
|
|
ledger, dashboard, and format-translation layer from scratch, in a
|
|
piece this repo would then own and maintain indefinitely, against a
|
|
target (LiteLLM) that already does all of it, is MIT-licensed, and is
|
|
a straightforward docker-compose service. Not adopted.
|
|
|
|
## Risks / gaps to carry into later tickets
|
|
|
|
1. **Lazytainer idle-suspend interaction (flagged in #9's "Not yet
|
|
specified").** Lazytainer decides to stop `llama-server` based on
|
|
network packet activity on its published port
|
|
(`lazytainer.group.llamaserver.minPacketThreshold` /
|
|
`inactiveTimeout` in this repo's `docker-compose.yml`). If LiteLLM
|
|
proxy performs periodic background health checks against configured
|
|
models (a common gateway behavior), that traffic could look like
|
|
real usage to Lazytainer and prevent it from ever idling the
|
|
container down. This needs to be checked against LiteLLM's actual
|
|
health-check config (there are documented options to disable/tune
|
|
background health checks) once the compose service is authored in
|
|
#14, and verified on real hardware per #9's "Not yet specified" note
|
|
on Lazytainer + multi-workload proxy interaction.
|
|
2. **llama.cpp's native `/v1/messages` shim vs. LiteLLM's own Anthropic
|
|
translation.** Two wiring options exist (LiteLLM translates
|
|
Anthropic→OpenAI itself and hits llama.cpp's OpenAI endpoint; or
|
|
LiteLLM passes Anthropic-format requests straight through to
|
|
llama.cpp's own native shim). This ticket confirms both are
|
|
plausible per LiteLLM's docs but doesn't pick one — that's
|
|
implementation detail for #14/#15, and should be smoke-tested against
|
|
the actual coding-CLI flows in `docs/coding-cli-setup.md` once wired.
|
|
3. **Priority queueing for #16 is a real but beta/rough-edged LiteLLM
|
|
feature**, with at least one reported-and-declined bug in the exact
|
|
`priority` mechanism, and the stronger reserved-capacity variant is
|
|
enterprise-gated. #16 should treat LiteLLM's scheduler as a
|
|
starting point to validate hands-on, not an assumed solved problem —
|
|
if it doesn't hold up under test, #16 may need a lightweight queuing
|
|
shim in front of the proxy (e.g. a small request-queue sidecar) rather
|
|
than reworking the whole gateway choice.
|
|
4. **Redis need is deferred, not eliminated.** A single-instance LiteLLM
|
|
deployment (the right size for this effort) doesn't need Redis for
|
|
virtual keys/spend/dashboard, but the priority scheduler's
|
|
multi-instance behavior and rate-limit sharing do use it — if #16
|
|
ends up needing Redis-backed prioritization even in a single-instance
|
|
deployment, add a `redis` service to the compose file at that point;
|
|
no need to provision it speculatively now.
|
|
|
|
## Bottom line for the wayfinder map
|
|
|
|
- Adopt **LiteLLM proxy** (MIT-licensed, `BerriAI/litellm`) as the
|
|
gateway/proxy in front of llama.cpp for this effort.
|
|
- It meets every hard requirement in #10 in its free/OSS build: virtual
|
|
keys with per-key spend, a self-hosted Admin UI spend dashboard,
|
|
config-driven custom per-token pricing (feeds #11 directly), a
|
|
documented two-service docker-compose deployment, and an
|
|
arbitrary-provider `model_list` that keeps future backends a config
|
|
change away.
|
|
- Its beta priority/queueing feature is a promising but unproven fit for
|
|
#16 — validate it hands-on rather than assuming it's settled.
|
|
- Portkey's self-hosted OSS gateway is disqualified (no self-hosted
|
|
dashboard/budgets). Helicone is a workable but weaker fallback
|
|
(maintenance-mode signal, feature-reduced self-host build, less clearly
|
|
documented virtual-key/custom-pricing support) if LiteLLM turns out to
|
|
be a poor fit during implementation.
|