211 lines
12 KiB
Markdown
211 lines
12 KiB
Markdown
# Research: Reference cloud model/pricing for the shadow-cost estimate
|
|
|
|
**Question:** Which reference cloud model/API pricing should the shadow-cost
|
|
estimate use (per #9's Destination), and how does that rate get wired into
|
|
LiteLLM (per #10's tool choice) to compute "what this local usage would have
|
|
cost on a real cloud API"?
|
|
|
|
**Answer:** Use a single fixed reference — **Claude Sonnet 5**, at its
|
|
current published API price, hardcoded into one `model_info` block in
|
|
LiteLLM's `config.yaml`. Skip a multi-tier pricing table.
|
|
|
|
## Reference model: single fixed price, not a tier table
|
|
|
|
Two options were weighed, per #11's framing:
|
|
|
|
**Option A — single fixed reference (recommended).** Pick one current Claude
|
|
model and price the local model against it always.
|
|
|
|
**Option B — small pricing table across a couple of tiers** (e.g. price the
|
|
same local usage against both a Haiku-tier and a Sonnet-tier rate
|
|
simultaneously, showing a range).
|
|
|
|
Recommendation: **Option A**, using **Claude Sonnet 5** — the model this
|
|
Claude Code CLI session itself runs on, and the model this repo's own
|
|
`docs/coding-cli-setup.md` documents pointing coding CLIs at this stack's
|
|
local Qwen3.8-27B model. Reasoning:
|
|
|
|
- **Matches what the user already tracks.** Map #9's Notes says the user
|
|
already tracks Claude pricing day-to-day. Sonnet is the tier actually used
|
|
for coding-CLI work against Claude directly (and is what this session is
|
|
running as), not a hypothetical comparison tier — a single number tied to
|
|
"what I'd have paid on the tool I actually use" is more meaningful than an
|
|
abstract low/high band.
|
|
- **Simple to maintain.** One `model_info` block, one number to update when
|
|
Anthropic changes Sonnet pricing (rare — pricing on this page has been
|
|
stable for months at a time; see staleness section below), versus a table
|
|
that needs every row kept current. This is explicitly "for fun" (map #9's
|
|
Destination/Notes), not real accounting — a table adds bookkeeping
|
|
overhead the use case doesn't need.
|
|
- **Local model's size is Sonnet-comparable, not flagship-comparable.** The
|
|
locally hosted model is `Qwen3.8-27B` (`docs/coding-cli-setup.md`) — a
|
|
~27B-class model. Pricing it against Anthropic's flagship (Opus tier,
|
|
$5/$25 per MTok) would overstate the shadow cost for what's actually a
|
|
mid-tier local model; pricing it against Sonnet ($2/$10 per MTok, the
|
|
mid-tier) is the closer-fitting comparison, and it's also the tier this
|
|
repo already documents pointing coding CLIs at when using the *real*
|
|
Claude API (as opposed to the local shim) is desired.
|
|
- A tier table would matter if the goal were "estimate real cloud spend
|
|
across scenarios," but map #9 explicitly frames this as a shadow/for-fun
|
|
estimate with no real external routing wired in — one clear number serves
|
|
that better than a range that needs interpreting.
|
|
|
|
**Current price** (as of 2026-08-25, per Anthropic's official pricing page):
|
|
|
|
| Model | Input | Output |
|
|
|---|---|---|
|
|
| **Claude Sonnet 5** | **$2.00 / MTok** ($0.000002/token) | **$10.00 / MTok** ($0.00001/token) |
|
|
|
|
Source: [Claude Platform docs — Pricing](https://platform.claude.com/docs/en/about-claude/pricing)
|
|
(Model pricing table). Note: this $2/$10 rate was originally introductory
|
|
pricing through 2026-08-31 with a scheduled increase to $3/$15 on
|
|
2026-09-01; Anthropic's pricing page (fetched today) states that increase
|
|
"will not occur" and $2/$10 is now the standard price — so this is a stable
|
|
number, not a rate about to change out from under the config.
|
|
|
|
Ignore prompt-caching, batch, and tool-use-overhead pricing modifiers for
|
|
this shadow estimate — llama.cpp's local usage has none of those API
|
|
features, so there's nothing to map them onto; base input/output token
|
|
pricing is the only thing that has a clean local-usage analogue.
|
|
|
|
## LiteLLM config shape: `model_info` custom pricing
|
|
|
|
Confirmed against LiteLLM's docs (same mechanism #10 already found; this
|
|
plugs the confirmed rate straight in). Add a `model_info` block to the local
|
|
model's entry in `config.yaml`:
|
|
|
|
```yaml
|
|
model_list:
|
|
- model_name: qwen3.8-27b-local # or whatever this stack names it
|
|
litellm_params:
|
|
model: openai/qwen3.8-27b # or the provider shim used to reach llama.cpp
|
|
api_base: http://llama-server:8080/v1
|
|
model_info:
|
|
input_cost_per_token: 0.000002 # $2 / 1,000,000 — Claude Sonnet 5 input rate
|
|
output_cost_per_token: 0.00001 # $10 / 1,000,000 — Claude Sonnet 5 output rate
|
|
```
|
|
|
|
Confirmed details:
|
|
|
|
- **Exact keys**: `model_info.input_cost_per_token` and
|
|
`model_info.output_cost_per_token`, both plain decimal USD-per-token
|
|
floats. LiteLLM also supports `input_cost_per_second` (time-based, e.g.
|
|
SageMaker-style billing) and `input_cost_per_character` /
|
|
`input_cost_per_image` / `input_cost_per_audio_token` /
|
|
`input_cost_per_video_per_second` for other modalities — none needed here
|
|
since this is a plain text chat model priced token-for-token.
|
|
- **Input vs. output distinguished**: yes — separate keys, matching how
|
|
Claude's own pricing (and llama.cpp's own `usage.prompt_tokens` /
|
|
`usage.completion_tokens` split) is already input/output-separated.
|
|
- **Per-model override**: yes — `model_info` is set per entry in
|
|
`model_list`, so only the local model entry needs it; LiteLLM's own
|
|
built-in cost map for 100+ known providers is untouched for any other
|
|
model added to the proxy later.
|
|
- **Where the resulting cost surfaces**: computed via LiteLLM's internal
|
|
`completion_cost()` function (the same path used for every provider,
|
|
built-in or custom-priced) on every `/chat/completions` /
|
|
`/v1/messages` call. It surfaces in:
|
|
- **Per-key spend**: `/key/info` returns a `spend` field (cumulative USD)
|
|
per virtual key — this is the per-workload number map #9 wants.
|
|
- **Admin UI dashboard** (`/ui`, confirmed present in #10's research):
|
|
the Usage tab visualizes spend by key/team, sourced from the same
|
|
spend ledger.
|
|
- **Spend logs**: written to LiteLLM's Postgres-backed
|
|
`LiteLLM_SpendLogs` / verification-token table, queryable via
|
|
`/team/info` and `/user/info` for aggregation.
|
|
- **Per-call logging object**: `kwargs["response_cost"]` on each
|
|
completion call, for anyone hooking custom logging/callbacks later.
|
|
|
|
Source: [LiteLLM — Custom Pricing docs](https://docs.litellm.ai/docs/proxy/custom_pricing),
|
|
[LiteLLM — Virtual Keys docs](https://docs.litellm.ai/docs/proxy/virtual_keys)
|
|
(`/key/info` spend field example), [LiteLLM — completion_cost /
|
|
model_prices_and_context_window.json](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json)
|
|
(the built-in cost map that `model_info` overrides for a given entry).
|
|
|
|
## Token-count mapping: token-for-token is fine here
|
|
|
|
Claude and Qwen3.8-27B use different tokenizers, so the *same text* produces
|
|
different token counts on each — a genuinely rigorous "what would this
|
|
exact conversation have cost on Claude" would need to re-tokenize the local
|
|
conversation text with Claude's tokenizer and price *that* count, not
|
|
llama.cpp's own token count.
|
|
|
|
LiteLLM does not do this re-tokenization for a custom-priced model: its
|
|
`completion_cost()` multiplies the token counts reported in the backend's
|
|
own `usage.prompt_tokens` / `usage.completion_tokens` response by whatever
|
|
`input_cost_per_token` / `output_cost_per_token` is configured for that
|
|
`model_list` entry — it does not re-count tokens against a different
|
|
provider's tokenizer for cost purposes (LiteLLM's own token-counting
|
|
utilities, e.g. `token_counter()`, exist as a fallback for providers that
|
|
don't return usage at all, not as a re-tokenization step for cost
|
|
calculation on a provider that does).
|
|
|
|
**This is fine, and doesn't need fixing**, given map #9's own framing: this
|
|
is explicitly a shadow/for-fun estimate, not real accounting. Treating
|
|
llama.cpp's reported token count as if it were "Claude tokens" and applying
|
|
Claude's per-token rate directly is a reasonable, cheap approximation —
|
|
token counts between modern tokenizers for English text are typically
|
|
within a similar order of magnitude (roughly comparable, not identical), so
|
|
the estimate is order-of-magnitude meaningful ("this conversation would
|
|
have cost about $X on Claude") without claiming precision it doesn't have.
|
|
Building actual re-tokenization against Claude's tokenizer purely to feed a
|
|
for-fun number would be effort disproportionate to the destination. If this
|
|
ever needs to be exact, the fix is a small conversion factor applied at
|
|
config time (e.g. inflate the configured per-token rate by a fudge factor
|
|
to roughly account for tokenizer differences) — not worth doing now.
|
|
|
|
One tokenizer wrinkle worth noting for future-proofing, not action:
|
|
Anthropic's pricing page notes Claude 4.7-and-later models (which includes
|
|
Sonnet 5) use a newer tokenizer producing "approximately 30% more tokens for
|
|
the same text" than earlier Claude models. This doesn't change the
|
|
recommendation (Sonnet 5 is still the right reference), it's just a reminder
|
|
that "tokens" are already an approximate, provider-specific unit even within
|
|
Anthropic's own model lineup — reinforcing that treating llama.cpp's token
|
|
count as directly billable at Claude's rate is consistent with how loosely
|
|
"a token" is already defined across models, not a special-case shortcut
|
|
being taken here.
|
|
|
|
## Config staleness risk
|
|
|
|
LiteLLM ships a built-in default pricing table
|
|
([`model_prices_and_context_window.json`](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json))
|
|
covering 100+ known providers/models, refreshed via the LiteLLM project's
|
|
own releases — that table is what could silently drift out of date for any
|
|
model relying on it. **This doesn't apply to the local model here**: because
|
|
the local llama.cpp model isn't a real named provider model, its pricing is
|
|
only ever set via the explicit `model_info` block in this repo's own
|
|
`config.yaml`, which LiteLLM never overwrites or auto-refreshes — the only
|
|
way the shadow-cost number goes stale is if Anthropic changes Sonnet 5's
|
|
published price and nobody updates the two numbers in this repo's config.
|
|
|
|
Given Anthropic's pricing page explicitly states the Sonnet 5 rate is now
|
|
locked in as the standard price (the previously-scheduled 2026-09-01
|
|
increase was cancelled), staleness risk here is low and infrequent — but
|
|
not zero, since Anthropic can still change prices with future model
|
|
launches or repricing. Practical mitigation for implementation (#14): put
|
|
the reference price in `config.yaml` with a comment noting the source URL
|
|
and date it was last checked, so a future price change is a one-line
|
|
`config.yaml` edit plus a comment-date bump — no code change, no migration.
|
|
No need for anything more automated (e.g. scraping Anthropic's pricing page
|
|
at startup) — that's more machinery than a for-fun estimate warrants.
|
|
|
|
## Bottom line for #14 (compose/config authoring)
|
|
|
|
- Use **Claude Sonnet 5** as the fixed shadow-pricing reference:
|
|
`input_cost_per_token: 0.000002`, `output_cost_per_token: 0.00001` in the
|
|
local model's `model_info` block in LiteLLM's `config.yaml`.
|
|
- No pricing table, no per-request tier selection — one rate, one config
|
|
block, matches map #9's "for fun not real accounting" framing and the
|
|
user's existing day-to-day use of Claude Sonnet for coding-CLI work.
|
|
- Token counts come straight from llama.cpp's own reported
|
|
`usage.prompt_tokens`/`completion_tokens` via LiteLLM's normal
|
|
`completion_cost()` path — no re-tokenization against Claude's tokenizer,
|
|
which is an acceptable approximation for a shadow estimate.
|
|
- Resulting spend is visible per-key via `/key/info`, aggregated via
|
|
`/team/info`/`/user/info`, and in the Admin UI's Usage dashboard — no
|
|
extra tracking code needed, this is the same spend-tracking path LiteLLM
|
|
uses for any other provider.
|
|
- Staleness risk is low (Sonnet 5's rate is currently locked as standard
|
|
pricing) and, if it ever drifts, is a one-line `config.yaml` edit — worth
|
|
a source-URL-and-date comment in the config, nothing more elaborate.
|