From a751e656e3da6ee9c7ec5744e6ae679331036444 Mon Sep 17 00:00:00 2001 From: Haylan Date: Tue, 25 Aug 2026 06:49:38 +0200 Subject: [PATCH] research: reference cloud pricing for LiteLLM shadow-cost estimate Co-Authored-By: Claude Sonnet 5 --- docs/research/proxy-shadow-pricing.md | 210 ++++++++++++++++++++++++++ 1 file changed, 210 insertions(+) create mode 100644 docs/research/proxy-shadow-pricing.md diff --git a/docs/research/proxy-shadow-pricing.md b/docs/research/proxy-shadow-pricing.md new file mode 100644 index 0000000..6f54330 --- /dev/null +++ b/docs/research/proxy-shadow-pricing.md @@ -0,0 +1,210 @@ +# Research: Reference cloud model/pricing for the shadow-cost estimate + +**Question:** Which reference cloud model/API pricing should the shadow-cost +estimate use (per #9's Destination), and how does that rate get wired into +LiteLLM (per #10's tool choice) to compute "what this local usage would have +cost on a real cloud API"? + +**Answer:** Use a single fixed reference — **Claude Sonnet 5**, at its +current published API price, hardcoded into one `model_info` block in +LiteLLM's `config.yaml`. Skip a multi-tier pricing table. + +## Reference model: single fixed price, not a tier table + +Two options were weighed, per #11's framing: + +**Option A — single fixed reference (recommended).** Pick one current Claude +model and price the local model against it always. + +**Option B — small pricing table across a couple of tiers** (e.g. price the +same local usage against both a Haiku-tier and a Sonnet-tier rate +simultaneously, showing a range). + +Recommendation: **Option A**, using **Claude Sonnet 5** — the model this +Claude Code CLI session itself runs on, and the model this repo's own +`docs/coding-cli-setup.md` documents pointing coding CLIs at this stack's +local Qwen3.8-27B model. Reasoning: + +- **Matches what the user already tracks.** Map #9's Notes says the user + already tracks Claude pricing day-to-day. Sonnet is the tier actually used + for coding-CLI work against Claude directly (and is what this session is + running as), not a hypothetical comparison tier — a single number tied to + "what I'd have paid on the tool I actually use" is more meaningful than an + abstract low/high band. +- **Simple to maintain.** One `model_info` block, one number to update when + Anthropic changes Sonnet pricing (rare — pricing on this page has been + stable for months at a time; see staleness section below), versus a table + that needs every row kept current. This is explicitly "for fun" (map #9's + Destination/Notes), not real accounting — a table adds bookkeeping + overhead the use case doesn't need. +- **Local model's size is Sonnet-comparable, not flagship-comparable.** The + locally hosted model is `Qwen3.8-27B` (`docs/coding-cli-setup.md`) — a + ~27B-class model. Pricing it against Anthropic's flagship (Opus tier, + $5/$25 per MTok) would overstate the shadow cost for what's actually a + mid-tier local model; pricing it against Sonnet ($2/$10 per MTok, the + mid-tier) is the closer-fitting comparison, and it's also the tier this + repo already documents pointing coding CLIs at when using the *real* + Claude API (as opposed to the local shim) is desired. +- A tier table would matter if the goal were "estimate real cloud spend + across scenarios," but map #9 explicitly frames this as a shadow/for-fun + estimate with no real external routing wired in — one clear number serves + that better than a range that needs interpreting. + +**Current price** (as of 2026-08-25, per Anthropic's official pricing page): + +| Model | Input | Output | +|---|---|---| +| **Claude Sonnet 5** | **$2.00 / MTok** ($0.000002/token) | **$10.00 / MTok** ($0.00001/token) | + +Source: [Claude Platform docs — Pricing](https://platform.claude.com/docs/en/about-claude/pricing) +(Model pricing table). Note: this $2/$10 rate was originally introductory +pricing through 2026-08-31 with a scheduled increase to $3/$15 on +2026-09-01; Anthropic's pricing page (fetched today) states that increase +"will not occur" and $2/$10 is now the standard price — so this is a stable +number, not a rate about to change out from under the config. + +Ignore prompt-caching, batch, and tool-use-overhead pricing modifiers for +this shadow estimate — llama.cpp's local usage has none of those API +features, so there's nothing to map them onto; base input/output token +pricing is the only thing that has a clean local-usage analogue. + +## LiteLLM config shape: `model_info` custom pricing + +Confirmed against LiteLLM's docs (same mechanism #10 already found; this +plugs the confirmed rate straight in). Add a `model_info` block to the local +model's entry in `config.yaml`: + +```yaml +model_list: + - model_name: qwen3.8-27b-local # or whatever this stack names it + litellm_params: + model: openai/qwen3.8-27b # or the provider shim used to reach llama.cpp + api_base: http://llama-server:8080/v1 + model_info: + input_cost_per_token: 0.000002 # $2 / 1,000,000 — Claude Sonnet 5 input rate + output_cost_per_token: 0.00001 # $10 / 1,000,000 — Claude Sonnet 5 output rate +``` + +Confirmed details: + +- **Exact keys**: `model_info.input_cost_per_token` and + `model_info.output_cost_per_token`, both plain decimal USD-per-token + floats. LiteLLM also supports `input_cost_per_second` (time-based, e.g. + SageMaker-style billing) and `input_cost_per_character` / + `input_cost_per_image` / `input_cost_per_audio_token` / + `input_cost_per_video_per_second` for other modalities — none needed here + since this is a plain text chat model priced token-for-token. +- **Input vs. output distinguished**: yes — separate keys, matching how + Claude's own pricing (and llama.cpp's own `usage.prompt_tokens` / + `usage.completion_tokens` split) is already input/output-separated. +- **Per-model override**: yes — `model_info` is set per entry in + `model_list`, so only the local model entry needs it; LiteLLM's own + built-in cost map for 100+ known providers is untouched for any other + model added to the proxy later. +- **Where the resulting cost surfaces**: computed via LiteLLM's internal + `completion_cost()` function (the same path used for every provider, + built-in or custom-priced) on every `/chat/completions` / + `/v1/messages` call. It surfaces in: + - **Per-key spend**: `/key/info` returns a `spend` field (cumulative USD) + per virtual key — this is the per-workload number map #9 wants. + - **Admin UI dashboard** (`/ui`, confirmed present in #10's research): + the Usage tab visualizes spend by key/team, sourced from the same + spend ledger. + - **Spend logs**: written to LiteLLM's Postgres-backed + `LiteLLM_SpendLogs` / verification-token table, queryable via + `/team/info` and `/user/info` for aggregation. + - **Per-call logging object**: `kwargs["response_cost"]` on each + completion call, for anyone hooking custom logging/callbacks later. + +Source: [LiteLLM — Custom Pricing docs](https://docs.litellm.ai/docs/proxy/custom_pricing), +[LiteLLM — Virtual Keys docs](https://docs.litellm.ai/docs/proxy/virtual_keys) +(`/key/info` spend field example), [LiteLLM — completion_cost / +model_prices_and_context_window.json](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json) +(the built-in cost map that `model_info` overrides for a given entry). + +## Token-count mapping: token-for-token is fine here + +Claude and Qwen3.8-27B use different tokenizers, so the *same text* produces +different token counts on each — a genuinely rigorous "what would this +exact conversation have cost on Claude" would need to re-tokenize the local +conversation text with Claude's tokenizer and price *that* count, not +llama.cpp's own token count. + +LiteLLM does not do this re-tokenization for a custom-priced model: its +`completion_cost()` multiplies the token counts reported in the backend's +own `usage.prompt_tokens` / `usage.completion_tokens` response by whatever +`input_cost_per_token` / `output_cost_per_token` is configured for that +`model_list` entry — it does not re-count tokens against a different +provider's tokenizer for cost purposes (LiteLLM's own token-counting +utilities, e.g. `token_counter()`, exist as a fallback for providers that +don't return usage at all, not as a re-tokenization step for cost +calculation on a provider that does). + +**This is fine, and doesn't need fixing**, given map #9's own framing: this +is explicitly a shadow/for-fun estimate, not real accounting. Treating +llama.cpp's reported token count as if it were "Claude tokens" and applying +Claude's per-token rate directly is a reasonable, cheap approximation — +token counts between modern tokenizers for English text are typically +within a similar order of magnitude (roughly comparable, not identical), so +the estimate is order-of-magnitude meaningful ("this conversation would +have cost about $X on Claude") without claiming precision it doesn't have. +Building actual re-tokenization against Claude's tokenizer purely to feed a +for-fun number would be effort disproportionate to the destination. If this +ever needs to be exact, the fix is a small conversion factor applied at +config time (e.g. inflate the configured per-token rate by a fudge factor +to roughly account for tokenizer differences) — not worth doing now. + +One tokenizer wrinkle worth noting for future-proofing, not action: +Anthropic's pricing page notes Claude 4.7-and-later models (which includes +Sonnet 5) use a newer tokenizer producing "approximately 30% more tokens for +the same text" than earlier Claude models. This doesn't change the +recommendation (Sonnet 5 is still the right reference), it's just a reminder +that "tokens" are already an approximate, provider-specific unit even within +Anthropic's own model lineup — reinforcing that treating llama.cpp's token +count as directly billable at Claude's rate is consistent with how loosely +"a token" is already defined across models, not a special-case shortcut +being taken here. + +## Config staleness risk + +LiteLLM ships a built-in default pricing table +([`model_prices_and_context_window.json`](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json)) +covering 100+ known providers/models, refreshed via the LiteLLM project's +own releases — that table is what could silently drift out of date for any +model relying on it. **This doesn't apply to the local model here**: because +the local llama.cpp model isn't a real named provider model, its pricing is +only ever set via the explicit `model_info` block in this repo's own +`config.yaml`, which LiteLLM never overwrites or auto-refreshes — the only +way the shadow-cost number goes stale is if Anthropic changes Sonnet 5's +published price and nobody updates the two numbers in this repo's config. + +Given Anthropic's pricing page explicitly states the Sonnet 5 rate is now +locked in as the standard price (the previously-scheduled 2026-09-01 +increase was cancelled), staleness risk here is low and infrequent — but +not zero, since Anthropic can still change prices with future model +launches or repricing. Practical mitigation for implementation (#14): put +the reference price in `config.yaml` with a comment noting the source URL +and date it was last checked, so a future price change is a one-line +`config.yaml` edit plus a comment-date bump — no code change, no migration. +No need for anything more automated (e.g. scraping Anthropic's pricing page +at startup) — that's more machinery than a for-fun estimate warrants. + +## Bottom line for #14 (compose/config authoring) + +- Use **Claude Sonnet 5** as the fixed shadow-pricing reference: + `input_cost_per_token: 0.000002`, `output_cost_per_token: 0.00001` in the + local model's `model_info` block in LiteLLM's `config.yaml`. +- No pricing table, no per-request tier selection — one rate, one config + block, matches map #9's "for fun not real accounting" framing and the + user's existing day-to-day use of Claude Sonnet for coding-CLI work. +- Token counts come straight from llama.cpp's own reported + `usage.prompt_tokens`/`completion_tokens` via LiteLLM's normal + `completion_cost()` path — no re-tokenization against Claude's tokenizer, + which is an acceptable approximation for a shadow estimate. +- Resulting spend is visible per-key via `/key/info`, aggregated via + `/team/info`/`/user/info`, and in the Admin UI's Usage dashboard — no + extra tracking code needed, this is the same spend-tracking path LiteLLM + uses for any other provider. +- Staleness risk is low (Sonnet 5's rate is currently locked as standard + pricing) and, if it ever drifts, is a one-line `config.yaml` edit — worth + a source-URL-and-date comment in the config, nothing more elaborate.