Files
LLM-Server/docs/research/proxy-shadow-pricing.md

12 KiB

Research: Reference cloud model/pricing for the shadow-cost estimate

Question: Which reference cloud model/API pricing should the shadow-cost estimate use (per #9's Destination), and how does that rate get wired into LiteLLM (per #10's tool choice) to compute "what this local usage would have cost on a real cloud API"?

Answer: Use a single fixed reference — Claude Sonnet 5, at its current published API price, hardcoded into one model_info block in LiteLLM's config.yaml. Skip a multi-tier pricing table.

Reference model: single fixed price, not a tier table

Two options were weighed, per #11's framing:

Option A — single fixed reference (recommended). Pick one current Claude model and price the local model against it always.

Option B — small pricing table across a couple of tiers (e.g. price the same local usage against both a Haiku-tier and a Sonnet-tier rate simultaneously, showing a range).

Recommendation: Option A, using Claude Sonnet 5 — the model this Claude Code CLI session itself runs on, and the model this repo's own docs/coding-cli-setup.md documents pointing coding CLIs at this stack's local Qwen3.8-27B model. Reasoning:

  • Matches what the user already tracks. Map #9's Notes says the user already tracks Claude pricing day-to-day. Sonnet is the tier actually used for coding-CLI work against Claude directly (and is what this session is running as), not a hypothetical comparison tier — a single number tied to "what I'd have paid on the tool I actually use" is more meaningful than an abstract low/high band.
  • Simple to maintain. One model_info block, one number to update when Anthropic changes Sonnet pricing (rare — pricing on this page has been stable for months at a time; see staleness section below), versus a table that needs every row kept current. This is explicitly "for fun" (map #9's Destination/Notes), not real accounting — a table adds bookkeeping overhead the use case doesn't need.
  • Local model's size is Sonnet-comparable, not flagship-comparable. The locally hosted model is Qwen3.8-27B (docs/coding-cli-setup.md) — a ~27B-class model. Pricing it against Anthropic's flagship (Opus tier, $5/$25 per MTok) would overstate the shadow cost for what's actually a mid-tier local model; pricing it against Sonnet ($2/$10 per MTok, the mid-tier) is the closer-fitting comparison, and it's also the tier this repo already documents pointing coding CLIs at when using the real Claude API (as opposed to the local shim) is desired.
  • A tier table would matter if the goal were "estimate real cloud spend across scenarios," but map #9 explicitly frames this as a shadow/for-fun estimate with no real external routing wired in — one clear number serves that better than a range that needs interpreting.

Current price (as of 2026-08-25, per Anthropic's official pricing page):

Model Input Output
Claude Sonnet 5 $2.00 / MTok ($0.000002/token) $10.00 / MTok ($0.00001/token)

Source: Claude Platform docs — Pricing (Model pricing table). Note: this $2/$10 rate was originally introductory pricing through 2026-08-31 with a scheduled increase to $3/$15 on 2026-09-01; Anthropic's pricing page (fetched today) states that increase "will not occur" and $2/$10 is now the standard price — so this is a stable number, not a rate about to change out from under the config.

Ignore prompt-caching, batch, and tool-use-overhead pricing modifiers for this shadow estimate — llama.cpp's local usage has none of those API features, so there's nothing to map them onto; base input/output token pricing is the only thing that has a clean local-usage analogue.

LiteLLM config shape: model_info custom pricing

Confirmed against LiteLLM's docs (same mechanism #10 already found; this plugs the confirmed rate straight in). Add a model_info block to the local model's entry in config.yaml:

model_list:
  - model_name: qwen3.8-27b-local          # or whatever this stack names it
    litellm_params:
      model: openai/qwen3.8-27b            # or the provider shim used to reach llama.cpp
      api_base: http://llama-server:8080/v1
    model_info:
      input_cost_per_token: 0.000002       # $2 / 1,000,000 — Claude Sonnet 5 input rate
      output_cost_per_token: 0.00001       # $10 / 1,000,000 — Claude Sonnet 5 output rate

Confirmed details:

  • Exact keys: model_info.input_cost_per_token and model_info.output_cost_per_token, both plain decimal USD-per-token floats. LiteLLM also supports input_cost_per_second (time-based, e.g. SageMaker-style billing) and input_cost_per_character / input_cost_per_image / input_cost_per_audio_token / input_cost_per_video_per_second for other modalities — none needed here since this is a plain text chat model priced token-for-token.
  • Input vs. output distinguished: yes — separate keys, matching how Claude's own pricing (and llama.cpp's own usage.prompt_tokens / usage.completion_tokens split) is already input/output-separated.
  • Per-model override: yes — model_info is set per entry in model_list, so only the local model entry needs it; LiteLLM's own built-in cost map for 100+ known providers is untouched for any other model added to the proxy later.
  • Where the resulting cost surfaces: computed via LiteLLM's internal completion_cost() function (the same path used for every provider, built-in or custom-priced) on every /chat/completions / /v1/messages call. It surfaces in:
    • Per-key spend: /key/info returns a spend field (cumulative USD) per virtual key — this is the per-workload number map #9 wants.
    • Admin UI dashboard (/ui, confirmed present in #10's research): the Usage tab visualizes spend by key/team, sourced from the same spend ledger.
    • Spend logs: written to LiteLLM's Postgres-backed LiteLLM_SpendLogs / verification-token table, queryable via /team/info and /user/info for aggregation.
    • Per-call logging object: kwargs["response_cost"] on each completion call, for anyone hooking custom logging/callbacks later.

Source: LiteLLM — Custom Pricing docs, LiteLLM — Virtual Keys docs (/key/info spend field example), LiteLLM — completion_cost / model_prices_and_context_window.json (the built-in cost map that model_info overrides for a given entry).

Token-count mapping: token-for-token is fine here

Claude and Qwen3.8-27B use different tokenizers, so the same text produces different token counts on each — a genuinely rigorous "what would this exact conversation have cost on Claude" would need to re-tokenize the local conversation text with Claude's tokenizer and price that count, not llama.cpp's own token count.

LiteLLM does not do this re-tokenization for a custom-priced model: its completion_cost() multiplies the token counts reported in the backend's own usage.prompt_tokens / usage.completion_tokens response by whatever input_cost_per_token / output_cost_per_token is configured for that model_list entry — it does not re-count tokens against a different provider's tokenizer for cost purposes (LiteLLM's own token-counting utilities, e.g. token_counter(), exist as a fallback for providers that don't return usage at all, not as a re-tokenization step for cost calculation on a provider that does).

This is fine, and doesn't need fixing, given map #9's own framing: this is explicitly a shadow/for-fun estimate, not real accounting. Treating llama.cpp's reported token count as if it were "Claude tokens" and applying Claude's per-token rate directly is a reasonable, cheap approximation — token counts between modern tokenizers for English text are typically within a similar order of magnitude (roughly comparable, not identical), so the estimate is order-of-magnitude meaningful ("this conversation would have cost about $X on Claude") without claiming precision it doesn't have. Building actual re-tokenization against Claude's tokenizer purely to feed a for-fun number would be effort disproportionate to the destination. If this ever needs to be exact, the fix is a small conversion factor applied at config time (e.g. inflate the configured per-token rate by a fudge factor to roughly account for tokenizer differences) — not worth doing now.

One tokenizer wrinkle worth noting for future-proofing, not action: Anthropic's pricing page notes Claude 4.7-and-later models (which includes Sonnet 5) use a newer tokenizer producing "approximately 30% more tokens for the same text" than earlier Claude models. This doesn't change the recommendation (Sonnet 5 is still the right reference), it's just a reminder that "tokens" are already an approximate, provider-specific unit even within Anthropic's own model lineup — reinforcing that treating llama.cpp's token count as directly billable at Claude's rate is consistent with how loosely "a token" is already defined across models, not a special-case shortcut being taken here.

Config staleness risk

LiteLLM ships a built-in default pricing table (model_prices_and_context_window.json) covering 100+ known providers/models, refreshed via the LiteLLM project's own releases — that table is what could silently drift out of date for any model relying on it. This doesn't apply to the local model here: because the local llama.cpp model isn't a real named provider model, its pricing is only ever set via the explicit model_info block in this repo's own config.yaml, which LiteLLM never overwrites or auto-refreshes — the only way the shadow-cost number goes stale is if Anthropic changes Sonnet 5's published price and nobody updates the two numbers in this repo's config.

Given Anthropic's pricing page explicitly states the Sonnet 5 rate is now locked in as the standard price (the previously-scheduled 2026-09-01 increase was cancelled), staleness risk here is low and infrequent — but not zero, since Anthropic can still change prices with future model launches or repricing. Practical mitigation for implementation (#14): put the reference price in config.yaml with a comment noting the source URL and date it was last checked, so a future price change is a one-line config.yaml edit plus a comment-date bump — no code change, no migration. No need for anything more automated (e.g. scraping Anthropic's pricing page at startup) — that's more machinery than a for-fun estimate warrants.

Bottom line for #14 (compose/config authoring)

  • Use Claude Sonnet 5 as the fixed shadow-pricing reference: input_cost_per_token: 0.000002, output_cost_per_token: 0.00001 in the local model's model_info block in LiteLLM's config.yaml.
  • No pricing table, no per-request tier selection — one rate, one config block, matches map #9's "for fun not real accounting" framing and the user's existing day-to-day use of Claude Sonnet for coding-CLI work.
  • Token counts come straight from llama.cpp's own reported usage.prompt_tokens/completion_tokens via LiteLLM's normal completion_cost() path — no re-tokenization against Claude's tokenizer, which is an acceptable approximation for a shadow estimate.
  • Resulting spend is visible per-key via /key/info, aggregated via /team/info//user/info, and in the Admin UI's Usage dashboard — no extra tracking code needed, this is the same spend-tracking path LiteLLM uses for any other provider.
  • Staleness risk is low (Sonnet 5's rate is currently locked as standard pricing) and, if it ever drifts, is a one-line config.yaml edit — worth a source-URL-and-date comment in the config, nothing more elaborate.