Files
LLM-Server/docs/research/proxy-shadow-pricing.md
T

211 lines
12 KiB
Markdown

# Research: Reference cloud model/pricing for the shadow-cost estimate
**Question:** Which reference cloud model/API pricing should the shadow-cost
estimate use (per #9's Destination), and how does that rate get wired into
LiteLLM (per #10's tool choice) to compute "what this local usage would have
cost on a real cloud API"?
**Answer:** Use a single fixed reference — **Claude Sonnet 5**, at its
current published API price, hardcoded into one `model_info` block in
LiteLLM's `config.yaml`. Skip a multi-tier pricing table.
## Reference model: single fixed price, not a tier table
Two options were weighed, per #11's framing:
**Option A — single fixed reference (recommended).** Pick one current Claude
model and price the local model against it always.
**Option B — small pricing table across a couple of tiers** (e.g. price the
same local usage against both a Haiku-tier and a Sonnet-tier rate
simultaneously, showing a range).
Recommendation: **Option A**, using **Claude Sonnet 5** — the model this
Claude Code CLI session itself runs on, and the model this repo's own
`docs/coding-cli-setup.md` documents pointing coding CLIs at this stack's
local Qwen3.8-27B model. Reasoning:
- **Matches what the user already tracks.** Map #9's Notes says the user
already tracks Claude pricing day-to-day. Sonnet is the tier actually used
for coding-CLI work against Claude directly (and is what this session is
running as), not a hypothetical comparison tier — a single number tied to
"what I'd have paid on the tool I actually use" is more meaningful than an
abstract low/high band.
- **Simple to maintain.** One `model_info` block, one number to update when
Anthropic changes Sonnet pricing (rare — pricing on this page has been
stable for months at a time; see staleness section below), versus a table
that needs every row kept current. This is explicitly "for fun" (map #9's
Destination/Notes), not real accounting — a table adds bookkeeping
overhead the use case doesn't need.
- **Local model's size is Sonnet-comparable, not flagship-comparable.** The
locally hosted model is `Qwen3.8-27B` (`docs/coding-cli-setup.md`) — a
~27B-class model. Pricing it against Anthropic's flagship (Opus tier,
$5/$25 per MTok) would overstate the shadow cost for what's actually a
mid-tier local model; pricing it against Sonnet ($2/$10 per MTok, the
mid-tier) is the closer-fitting comparison, and it's also the tier this
repo already documents pointing coding CLIs at when using the *real*
Claude API (as opposed to the local shim) is desired.
- A tier table would matter if the goal were "estimate real cloud spend
across scenarios," but map #9 explicitly frames this as a shadow/for-fun
estimate with no real external routing wired in — one clear number serves
that better than a range that needs interpreting.
**Current price** (as of 2026-08-25, per Anthropic's official pricing page):
| Model | Input | Output |
|---|---|---|
| **Claude Sonnet 5** | **$2.00 / MTok** ($0.000002/token) | **$10.00 / MTok** ($0.00001/token) |
Source: [Claude Platform docs — Pricing](https://platform.claude.com/docs/en/about-claude/pricing)
(Model pricing table). Note: this $2/$10 rate was originally introductory
pricing through 2026-08-31 with a scheduled increase to $3/$15 on
2026-09-01; Anthropic's pricing page (fetched today) states that increase
"will not occur" and $2/$10 is now the standard price — so this is a stable
number, not a rate about to change out from under the config.
Ignore prompt-caching, batch, and tool-use-overhead pricing modifiers for
this shadow estimate — llama.cpp's local usage has none of those API
features, so there's nothing to map them onto; base input/output token
pricing is the only thing that has a clean local-usage analogue.
## LiteLLM config shape: `model_info` custom pricing
Confirmed against LiteLLM's docs (same mechanism #10 already found; this
plugs the confirmed rate straight in). Add a `model_info` block to the local
model's entry in `config.yaml`:
```yaml
model_list:
- model_name: qwen3.8-27b-local # or whatever this stack names it
litellm_params:
model: openai/qwen3.8-27b # or the provider shim used to reach llama.cpp
api_base: http://llama-server:8080/v1
model_info:
input_cost_per_token: 0.000002 # $2 / 1,000,000 — Claude Sonnet 5 input rate
output_cost_per_token: 0.00001 # $10 / 1,000,000 — Claude Sonnet 5 output rate
```
Confirmed details:
- **Exact keys**: `model_info.input_cost_per_token` and
`model_info.output_cost_per_token`, both plain decimal USD-per-token
floats. LiteLLM also supports `input_cost_per_second` (time-based, e.g.
SageMaker-style billing) and `input_cost_per_character` /
`input_cost_per_image` / `input_cost_per_audio_token` /
`input_cost_per_video_per_second` for other modalities — none needed here
since this is a plain text chat model priced token-for-token.
- **Input vs. output distinguished**: yes — separate keys, matching how
Claude's own pricing (and llama.cpp's own `usage.prompt_tokens` /
`usage.completion_tokens` split) is already input/output-separated.
- **Per-model override**: yes — `model_info` is set per entry in
`model_list`, so only the local model entry needs it; LiteLLM's own
built-in cost map for 100+ known providers is untouched for any other
model added to the proxy later.
- **Where the resulting cost surfaces**: computed via LiteLLM's internal
`completion_cost()` function (the same path used for every provider,
built-in or custom-priced) on every `/chat/completions` /
`/v1/messages` call. It surfaces in:
- **Per-key spend**: `/key/info` returns a `spend` field (cumulative USD)
per virtual key — this is the per-workload number map #9 wants.
- **Admin UI dashboard** (`/ui`, confirmed present in #10's research):
the Usage tab visualizes spend by key/team, sourced from the same
spend ledger.
- **Spend logs**: written to LiteLLM's Postgres-backed
`LiteLLM_SpendLogs` / verification-token table, queryable via
`/team/info` and `/user/info` for aggregation.
- **Per-call logging object**: `kwargs["response_cost"]` on each
completion call, for anyone hooking custom logging/callbacks later.
Source: [LiteLLM — Custom Pricing docs](https://docs.litellm.ai/docs/proxy/custom_pricing),
[LiteLLM — Virtual Keys docs](https://docs.litellm.ai/docs/proxy/virtual_keys)
(`/key/info` spend field example), [LiteLLM — completion_cost /
model_prices_and_context_window.json](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json)
(the built-in cost map that `model_info` overrides for a given entry).
## Token-count mapping: token-for-token is fine here
Claude and Qwen3.8-27B use different tokenizers, so the *same text* produces
different token counts on each — a genuinely rigorous "what would this
exact conversation have cost on Claude" would need to re-tokenize the local
conversation text with Claude's tokenizer and price *that* count, not
llama.cpp's own token count.
LiteLLM does not do this re-tokenization for a custom-priced model: its
`completion_cost()` multiplies the token counts reported in the backend's
own `usage.prompt_tokens` / `usage.completion_tokens` response by whatever
`input_cost_per_token` / `output_cost_per_token` is configured for that
`model_list` entry — it does not re-count tokens against a different
provider's tokenizer for cost purposes (LiteLLM's own token-counting
utilities, e.g. `token_counter()`, exist as a fallback for providers that
don't return usage at all, not as a re-tokenization step for cost
calculation on a provider that does).
**This is fine, and doesn't need fixing**, given map #9's own framing: this
is explicitly a shadow/for-fun estimate, not real accounting. Treating
llama.cpp's reported token count as if it were "Claude tokens" and applying
Claude's per-token rate directly is a reasonable, cheap approximation —
token counts between modern tokenizers for English text are typically
within a similar order of magnitude (roughly comparable, not identical), so
the estimate is order-of-magnitude meaningful ("this conversation would
have cost about $X on Claude") without claiming precision it doesn't have.
Building actual re-tokenization against Claude's tokenizer purely to feed a
for-fun number would be effort disproportionate to the destination. If this
ever needs to be exact, the fix is a small conversion factor applied at
config time (e.g. inflate the configured per-token rate by a fudge factor
to roughly account for tokenizer differences) — not worth doing now.
One tokenizer wrinkle worth noting for future-proofing, not action:
Anthropic's pricing page notes Claude 4.7-and-later models (which includes
Sonnet 5) use a newer tokenizer producing "approximately 30% more tokens for
the same text" than earlier Claude models. This doesn't change the
recommendation (Sonnet 5 is still the right reference), it's just a reminder
that "tokens" are already an approximate, provider-specific unit even within
Anthropic's own model lineup — reinforcing that treating llama.cpp's token
count as directly billable at Claude's rate is consistent with how loosely
"a token" is already defined across models, not a special-case shortcut
being taken here.
## Config staleness risk
LiteLLM ships a built-in default pricing table
([`model_prices_and_context_window.json`](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json))
covering 100+ known providers/models, refreshed via the LiteLLM project's
own releases — that table is what could silently drift out of date for any
model relying on it. **This doesn't apply to the local model here**: because
the local llama.cpp model isn't a real named provider model, its pricing is
only ever set via the explicit `model_info` block in this repo's own
`config.yaml`, which LiteLLM never overwrites or auto-refreshes — the only
way the shadow-cost number goes stale is if Anthropic changes Sonnet 5's
published price and nobody updates the two numbers in this repo's config.
Given Anthropic's pricing page explicitly states the Sonnet 5 rate is now
locked in as the standard price (the previously-scheduled 2026-09-01
increase was cancelled), staleness risk here is low and infrequent — but
not zero, since Anthropic can still change prices with future model
launches or repricing. Practical mitigation for implementation (#14): put
the reference price in `config.yaml` with a comment noting the source URL
and date it was last checked, so a future price change is a one-line
`config.yaml` edit plus a comment-date bump — no code change, no migration.
No need for anything more automated (e.g. scraping Anthropic's pricing page
at startup) — that's more machinery than a for-fun estimate warrants.
## Bottom line for #14 (compose/config authoring)
- Use **Claude Sonnet 5** as the fixed shadow-pricing reference:
`input_cost_per_token: 0.000002`, `output_cost_per_token: 0.00001` in the
local model's `model_info` block in LiteLLM's `config.yaml`.
- No pricing table, no per-request tier selection — one rate, one config
block, matches map #9's "for fun not real accounting" framing and the
user's existing day-to-day use of Claude Sonnet for coding-CLI work.
- Token counts come straight from llama.cpp's own reported
`usage.prompt_tokens`/`completion_tokens` via LiteLLM's normal
`completion_cost()` path — no re-tokenization against Claude's tokenizer,
which is an acceptable approximation for a shadow estimate.
- Resulting spend is visible per-key via `/key/info`, aggregated via
`/team/info`/`/user/info`, and in the Admin UI's Usage dashboard — no
extra tracking code needed, this is the same spend-tracking path LiteLLM
uses for any other provider.
- Staleness risk is low (Sonnet 5's rate is currently locked as standard
pricing) and, if it ever drifts, is a one-line `config.yaml` edit — worth
a source-URL-and-date comment in the config, nothing more elaborate.