12 KiB
Research: Reference cloud model/pricing for the shadow-cost estimate
Question: Which reference cloud model/API pricing should the shadow-cost estimate use (per #9's Destination), and how does that rate get wired into LiteLLM (per #10's tool choice) to compute "what this local usage would have cost on a real cloud API"?
Answer: Use a single fixed reference — Claude Sonnet 5, at its
current published API price, hardcoded into one model_info block in
LiteLLM's config.yaml. Skip a multi-tier pricing table.
Reference model: single fixed price, not a tier table
Two options were weighed, per #11's framing:
Option A — single fixed reference (recommended). Pick one current Claude model and price the local model against it always.
Option B — small pricing table across a couple of tiers (e.g. price the same local usage against both a Haiku-tier and a Sonnet-tier rate simultaneously, showing a range).
Recommendation: Option A, using Claude Sonnet 5 — the model this
Claude Code CLI session itself runs on, and the model this repo's own
docs/coding-cli-setup.md documents pointing coding CLIs at this stack's
local Qwen3.8-27B model. Reasoning:
- Matches what the user already tracks. Map #9's Notes says the user already tracks Claude pricing day-to-day. Sonnet is the tier actually used for coding-CLI work against Claude directly (and is what this session is running as), not a hypothetical comparison tier — a single number tied to "what I'd have paid on the tool I actually use" is more meaningful than an abstract low/high band.
- Simple to maintain. One
model_infoblock, one number to update when Anthropic changes Sonnet pricing (rare — pricing on this page has been stable for months at a time; see staleness section below), versus a table that needs every row kept current. This is explicitly "for fun" (map #9's Destination/Notes), not real accounting — a table adds bookkeeping overhead the use case doesn't need. - Local model's size is Sonnet-comparable, not flagship-comparable. The
locally hosted model is
Qwen3.8-27B(docs/coding-cli-setup.md) — a ~27B-class model. Pricing it against Anthropic's flagship (Opus tier, $5/$25 per MTok) would overstate the shadow cost for what's actually a mid-tier local model; pricing it against Sonnet ($2/$10 per MTok, the mid-tier) is the closer-fitting comparison, and it's also the tier this repo already documents pointing coding CLIs at when using the real Claude API (as opposed to the local shim) is desired. - A tier table would matter if the goal were "estimate real cloud spend across scenarios," but map #9 explicitly frames this as a shadow/for-fun estimate with no real external routing wired in — one clear number serves that better than a range that needs interpreting.
Current price (as of 2026-08-25, per Anthropic's official pricing page):
| Model | Input | Output |
|---|---|---|
| Claude Sonnet 5 | $2.00 / MTok ($0.000002/token) | $10.00 / MTok ($0.00001/token) |
Source: Claude Platform docs — Pricing (Model pricing table). Note: this $2/$10 rate was originally introductory pricing through 2026-08-31 with a scheduled increase to $3/$15 on 2026-09-01; Anthropic's pricing page (fetched today) states that increase "will not occur" and $2/$10 is now the standard price — so this is a stable number, not a rate about to change out from under the config.
Ignore prompt-caching, batch, and tool-use-overhead pricing modifiers for this shadow estimate — llama.cpp's local usage has none of those API features, so there's nothing to map them onto; base input/output token pricing is the only thing that has a clean local-usage analogue.
LiteLLM config shape: model_info custom pricing
Confirmed against LiteLLM's docs (same mechanism #10 already found; this
plugs the confirmed rate straight in). Add a model_info block to the local
model's entry in config.yaml:
model_list:
- model_name: qwen3.8-27b-local # or whatever this stack names it
litellm_params:
model: openai/qwen3.8-27b # or the provider shim used to reach llama.cpp
api_base: http://llama-server:8080/v1
model_info:
input_cost_per_token: 0.000002 # $2 / 1,000,000 — Claude Sonnet 5 input rate
output_cost_per_token: 0.00001 # $10 / 1,000,000 — Claude Sonnet 5 output rate
Confirmed details:
- Exact keys:
model_info.input_cost_per_tokenandmodel_info.output_cost_per_token, both plain decimal USD-per-token floats. LiteLLM also supportsinput_cost_per_second(time-based, e.g. SageMaker-style billing) andinput_cost_per_character/input_cost_per_image/input_cost_per_audio_token/input_cost_per_video_per_secondfor other modalities — none needed here since this is a plain text chat model priced token-for-token. - Input vs. output distinguished: yes — separate keys, matching how
Claude's own pricing (and llama.cpp's own
usage.prompt_tokens/usage.completion_tokenssplit) is already input/output-separated. - Per-model override: yes —
model_infois set per entry inmodel_list, so only the local model entry needs it; LiteLLM's own built-in cost map for 100+ known providers is untouched for any other model added to the proxy later. - Where the resulting cost surfaces: computed via LiteLLM's internal
completion_cost()function (the same path used for every provider, built-in or custom-priced) on every/chat/completions//v1/messagescall. It surfaces in:- Per-key spend:
/key/inforeturns aspendfield (cumulative USD) per virtual key — this is the per-workload number map #9 wants. - Admin UI dashboard (
/ui, confirmed present in #10's research): the Usage tab visualizes spend by key/team, sourced from the same spend ledger. - Spend logs: written to LiteLLM's Postgres-backed
LiteLLM_SpendLogs/ verification-token table, queryable via/team/infoand/user/infofor aggregation. - Per-call logging object:
kwargs["response_cost"]on each completion call, for anyone hooking custom logging/callbacks later.
- Per-key spend:
Source: LiteLLM — Custom Pricing docs,
LiteLLM — Virtual Keys docs
(/key/info spend field example), LiteLLM — completion_cost /
model_prices_and_context_window.json
(the built-in cost map that model_info overrides for a given entry).
Token-count mapping: token-for-token is fine here
Claude and Qwen3.8-27B use different tokenizers, so the same text produces different token counts on each — a genuinely rigorous "what would this exact conversation have cost on Claude" would need to re-tokenize the local conversation text with Claude's tokenizer and price that count, not llama.cpp's own token count.
LiteLLM does not do this re-tokenization for a custom-priced model: its
completion_cost() multiplies the token counts reported in the backend's
own usage.prompt_tokens / usage.completion_tokens response by whatever
input_cost_per_token / output_cost_per_token is configured for that
model_list entry — it does not re-count tokens against a different
provider's tokenizer for cost purposes (LiteLLM's own token-counting
utilities, e.g. token_counter(), exist as a fallback for providers that
don't return usage at all, not as a re-tokenization step for cost
calculation on a provider that does).
This is fine, and doesn't need fixing, given map #9's own framing: this is explicitly a shadow/for-fun estimate, not real accounting. Treating llama.cpp's reported token count as if it were "Claude tokens" and applying Claude's per-token rate directly is a reasonable, cheap approximation — token counts between modern tokenizers for English text are typically within a similar order of magnitude (roughly comparable, not identical), so the estimate is order-of-magnitude meaningful ("this conversation would have cost about $X on Claude") without claiming precision it doesn't have. Building actual re-tokenization against Claude's tokenizer purely to feed a for-fun number would be effort disproportionate to the destination. If this ever needs to be exact, the fix is a small conversion factor applied at config time (e.g. inflate the configured per-token rate by a fudge factor to roughly account for tokenizer differences) — not worth doing now.
One tokenizer wrinkle worth noting for future-proofing, not action: Anthropic's pricing page notes Claude 4.7-and-later models (which includes Sonnet 5) use a newer tokenizer producing "approximately 30% more tokens for the same text" than earlier Claude models. This doesn't change the recommendation (Sonnet 5 is still the right reference), it's just a reminder that "tokens" are already an approximate, provider-specific unit even within Anthropic's own model lineup — reinforcing that treating llama.cpp's token count as directly billable at Claude's rate is consistent with how loosely "a token" is already defined across models, not a special-case shortcut being taken here.
Config staleness risk
LiteLLM ships a built-in default pricing table
(model_prices_and_context_window.json)
covering 100+ known providers/models, refreshed via the LiteLLM project's
own releases — that table is what could silently drift out of date for any
model relying on it. This doesn't apply to the local model here: because
the local llama.cpp model isn't a real named provider model, its pricing is
only ever set via the explicit model_info block in this repo's own
config.yaml, which LiteLLM never overwrites or auto-refreshes — the only
way the shadow-cost number goes stale is if Anthropic changes Sonnet 5's
published price and nobody updates the two numbers in this repo's config.
Given Anthropic's pricing page explicitly states the Sonnet 5 rate is now
locked in as the standard price (the previously-scheduled 2026-09-01
increase was cancelled), staleness risk here is low and infrequent — but
not zero, since Anthropic can still change prices with future model
launches or repricing. Practical mitigation for implementation (#14): put
the reference price in config.yaml with a comment noting the source URL
and date it was last checked, so a future price change is a one-line
config.yaml edit plus a comment-date bump — no code change, no migration.
No need for anything more automated (e.g. scraping Anthropic's pricing page
at startup) — that's more machinery than a for-fun estimate warrants.
Bottom line for #14 (compose/config authoring)
- Use Claude Sonnet 5 as the fixed shadow-pricing reference:
input_cost_per_token: 0.000002,output_cost_per_token: 0.00001in the local model'smodel_infoblock in LiteLLM'sconfig.yaml. - No pricing table, no per-request tier selection — one rate, one config block, matches map #9's "for fun not real accounting" framing and the user's existing day-to-day use of Claude Sonnet for coding-CLI work.
- Token counts come straight from llama.cpp's own reported
usage.prompt_tokens/completion_tokensvia LiteLLM's normalcompletion_cost()path — no re-tokenization against Claude's tokenizer, which is an acceptable approximation for a shadow estimate. - Resulting spend is visible per-key via
/key/info, aggregated via/team/info//user/info, and in the Admin UI's Usage dashboard — no extra tracking code needed, this is the same spend-tracking path LiteLLM uses for any other provider. - Staleness risk is low (Sonnet 5's rate is currently locked as standard
pricing) and, if it ever drifts, is a one-line
config.yamledit — worth a source-URL-and-date comment in the config, nothing more elaborate.