738 lines
56 KiB
Markdown
738 lines
56 KiB
Markdown
# Research: hardware roadmap to 500k-token context × 2 parallel agents (1M stretch)
|
||
|
||
**Date:** 2026-09-11
|
||
|
||
**Question:** What VRAM does 500k-token context × 2 parallel llama-server slots (and a 1M-token
|
||
stretch goal) actually cost for the Qwen3 family, and what hardware roadmap gets there from the
|
||
current single-R9700 setup — given the user's stated plan to add an older (PCIe 4.0) Threadripper
|
||
for lane count, reuse existing RAM/PSU (~200W headroom / one spare 8-pin), mix in already-owned
|
||
NVIDIA cards (GTX 1080 8GB, RTX 2080 8GB, GT 710 1GB) for the classifier role, and price used GPUs
|
||
at roughly $20-30/GB VRAM?
|
||
|
||
**Answer, short version:** The two goals ("500k × 2 parallel" and "1M stretch") turn out to need
|
||
**the same total VRAM budget** — because of how llama-server's `--ctx-size` and `--parallel` interact
|
||
(§2), 500k × 2 slots and a single 1M-token slot both require setting `--ctx-size 1000000`. At
|
||
`q4_0`-quantized KV cache that's **~33 GB** (weights + KV) for Qwen3.8-27B, at `q8_0` it's **~48 GB**,
|
||
at fp16 it's **~79 GB** — before compute-buffer overhead. That does not fit on the current single
|
||
32GB R9700 at any KV precision, and comfortably fits on two 32GB-class cards only at `q8_0`/`q4_0`.
|
||
The user's $20-30/GB pricing intuition holds for last-gen used consumer cards (RTX 3060 12GB) but
|
||
**not** for RTX 3090 24GB (~$44/GB currently) or a second R9700 (~$41/GB, new — no used market yet
|
||
for a card released mid-2026). The stated Threadripper plan needs to specifically target the
|
||
**non-PRO Threadripper 3000 series on sTRX4** (64 lanes, PCIe 4.0) — older Threadripper on the
|
||
original TR4 socket (1000/2000 series) is PCIe 3.0 only, which doesn't match the user's own PCIe 4.0
|
||
requirement. The power budget (~200W / one spare 8-pin) is exhausted by a *single* mid-tier used GPU
|
||
addition — a PSU upgrade is not optional past the very first stage. See §7 for the roadmap.
|
||
|
||
---
|
||
|
||
## 1. Current state (from this repo)
|
||
|
||
From `docker-compose.yml` and `.env.example` at the repo root:
|
||
|
||
- **Main model:** `Qwen3.8-27B-UD-Q4_K_XL.gguf` (17.6 GB weights), `--ctx-size 262144`,
|
||
`--parallel 2`, `--flash-attn on`, `--cache-type-k q8_0 --cache-type-v q8_0`, `--n-gpu-layers 999`,
|
||
on one AMD Radeon AI PRO R9700 (32GB, ROCm/HIP, `gfx1201`).
|
||
- **Classifier ("fast") model:** `Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`, `--ctx-size 65536`,
|
||
`--parallel 1`, `--n-gpu-layers 28` (partial offload), `--cache-type-k/v q4_0`, its own container on
|
||
the *same* R9700, sharing VRAM with the main model — see
|
||
[`docker-compose.yml`](../../docker-compose.yml) lines ~58-99 and
|
||
[`fast-model-choice.md`](fast-model-choice.md).
|
||
- `.env.example` already documents the exact fact this research turns on: *"Each slot gets
|
||
`LLAMA_CTX_SIZE / LLAMA_PARALLEL` tokens of context"* — i.e. today's 262144 ctx-size ÷ 2 parallel
|
||
slots means each real request only gets **~131K tokens**, not the full 262144, confirmed in-repo
|
||
before any external source was checked.
|
||
- Prior research already worked out the KV-cache formula for this exact model
|
||
([`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md)) — this doc reuses and extends that math for the
|
||
500k/1M targets rather than re-deriving it.
|
||
|
||
`docs/server-planing.md` describes a **different, earlier plan**: a 4× AMD Radeon AI PRO R9700 rig
|
||
on a Gigabyte MZ32-AR0 (single-socket SP3/EPYC, 128 PCIe 4.0 lanes), fully AMD/ROCm. The user's plan
|
||
in this ticket is not that — it pivots toward an older **Threadripper** (SP3's sibling desktop-HEDT
|
||
socket family, not SP3 itself) and explicitly wants to mix in already-owned **NVIDIA** cards. These
|
||
two plans are **not the same build** and, per §6, ROCm and CUDA cards cannot share one llama.cpp
|
||
process — they can only coexist as separate containers on separate cards. Treat `server-planing.md`
|
||
as superseded context, not the active plan, unless the user says otherwise.
|
||
|
||
---
|
||
|
||
## 2. llama-server parallelism: does each slot get its own full `--ctx-size`, or is it divided?
|
||
|
||
**Divided.** This is the single fact that changes the whole budget by 2×, confirmed from three
|
||
independent primary sources:
|
||
|
||
1. **This repo's own `.env.example`** (quoted above) already documents it for the current deployment.
|
||
2. **llama.cpp's own server README**, fetched directly
|
||
([`tools/server/README.md`](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)):
|
||
`--ctx-size (-c)`: *"size of the prompt context (default: 0, 0 = loaded from model)"*;
|
||
`--parallel (-np)`: *"number of server slots (default: -1, -1 = auto)"* — the docs list these as
|
||
independent flags, but don't spell out the division themselves.
|
||
3. **A real user's server log**, quoted verbatim in
|
||
[ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681), is the actual proof:
|
||
running with `--ctx-size 327680 --parallel 6` produces `n_ctx = 327680`,
|
||
`n_ctx_per_seq = 54613` — i.e. `327680 / 6 ≈ 54613`. The reporter explicitly asked for a
|
||
`--ctx-size-per-seq`-style flag to *avoid* this division; no such flag exists as of the fetch date.
|
||
|
||
Practical consequence: **to get 500,000 usable tokens on each of 2 parallel slots, `--ctx-size` must
|
||
be set to 1,000,000, not 500,000.** The KV cache is sized off the *total* `--ctx-size`
|
||
(`--kv-unified`, on by default when slots are auto per the README's `-kvu` entry, uses one shared
|
||
pool sized to the full `n_ctx`) — so the VRAM cost of "500k × 2 parallel" and "one 1M-token slot"
|
||
is **identical**: both require `--ctx-size 1000000`. This is a genuinely useful finding for the
|
||
roadmap — reaching the 500k×2 target and the 1M stretch goal cost the same VRAM; the only difference
|
||
is `--parallel 1` vs `--parallel 2` at deploy time, a config change with zero extra hardware cost.
|
||
|
||
`--cache-type-k` / `--cache-type-v` accept `f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1`
|
||
(default `f16`), per the same README fetch. The repo already uses `q8_0` on the main model and `q4_0`
|
||
on the classifier, so both quantization tiers used in the math below are already-proven-working
|
||
configurations in this stack, not hypothetical flags.
|
||
|
||
---
|
||
|
||
## 3. KV-cache math per model
|
||
|
||
### Qwen3.8-27B (hybrid Gated-DeltaNet / attention)
|
||
|
||
Reusing the architecture params already pulled from
|
||
[`Qwen/Qwen3.8-27B/config.json`](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json) in
|
||
[`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), re-verified directly for this doc: `num_hidden_layers=64`,
|
||
`full_attention_interval=4` → **16 of 64 layers are standard KV-caching attention**, the other 48 are
|
||
Gated DeltaNet linear-attention layers with a small, context-length-*independent* recurrent state
|
||
(tens of MB total, negligible next to the attention KV cache — ignored below).
|
||
`num_key_value_heads=4` (GQA), `head_dim=256`. Native context `max_position_embeddings=262144`
|
||
(YaRN-extensible to 1M per the model card — **both the 500k and 1M targets exceed native context and
|
||
require RoPE/YaRN scaling**, which is a real quality caveat, not just a memory one — Qwen has not
|
||
published independent long-context quality benchmarks past native length that this research found).
|
||
|
||
Per-token KV cache, fp16, both K and V, across the 16 full-attention layers:
|
||
|
||
```
|
||
16 layers × 2 (K+V) × 4 kv_heads × 256 head_dim × 2 bytes = 64 KiB/token
|
||
```
|
||
|
||
| Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 |
|
||
|---|---|---|---|
|
||
| 262,144 (current) | ~16.0 GiB | ~8.0 GiB | ~4.0 GiB |
|
||
| 500,000 | ~30.5 GiB | ~15.3 GiB | ~7.6 GiB |
|
||
| **1,000,000 (500k×2, or 1M stretch)** | **~61.0 GiB** | **~30.5 GiB** | **~15.3 GiB** |
|
||
|
||
(`q8_0` is 8-bit vs. fp16's 16-bit → exactly half; `q4_0` is 4-bit → exactly quarter, per llama.cpp's
|
||
own cache-type byte widths.)
|
||
|
||
### Qwen3-4B-Instruct-2507 (plain GQA transformer, classifier role)
|
||
|
||
From [`Qwen/Qwen3-4B-Instruct-2507/config.json`](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json)
|
||
(already pulled in [`fast-model-choice.md`](fast-model-choice.md)): `num_hidden_layers=36` — every
|
||
layer is standard attention here (no hybrid split), `num_key_value_heads=8`, `head_dim=128`.
|
||
|
||
```
|
||
36 layers × 2 (K+V) × 8 kv_heads × 128 head_dim × 2 bytes = 144 KiB/token
|
||
```
|
||
|
||
The classifier's real transcript ceiling is ~40-50K tokens (qwen-code's own
|
||
`MAX_TRANSCRIPT_MESSAGES=40` × `MAX_HISTORICAL_ACTION_CHARS=4000`, per `fast-model-choice.md` §"what
|
||
actually shipped") — nowhere near 500k/1M, so the classifier does **not** need to grow for this
|
||
roadmap; it stays exactly as deployed today, on its own small allocation. Per-token cost is included
|
||
here only because it feeds the "does the classifier's dedicated GPU need to change" question in §6.
|
||
|
||
---
|
||
|
||
## 4. Total VRAM budget: 500k × 2 parallel, and the 1M stretch
|
||
|
||
Weights: `Qwen3.8-27B-UD-Q4_K_XL.gguf` is **17.6 GB**, confirmed directly from the
|
||
[unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main)
|
||
(already verified in `qwen3.8-27b-quant.md`).
|
||
|
||
Per §2, both "500k × 2 parallel" and "1M stretch" require `--ctx-size 1000000` — same KV budget:
|
||
|
||
| KV precision | KV cache | + weights (17.6 GB) | + est. compute-buffer/runtime overhead* | **Realistic total** |
|
||
|---|---|---|---|---|
|
||
| fp16 (default) | 61.0 GiB | 78.6 GiB | +3-6 GiB | **~82-85 GB** |
|
||
| q8_0 (proven in this stack today) | 30.5 GiB | 48.1 GiB | +3-6 GiB | **~51-54 GB** |
|
||
| q4_0 (proven in this stack today, on the classifier) | 15.3 GiB | 32.9 GiB | +3-6 GiB | **~36-39 GB** |
|
||
|
||
\* *Estimate, not a cited figure* — llama.cpp's flash-attention compute buffer scales closer to
|
||
linear than the unfused-attention path, per this repo's own measured note in `docker-compose.yml`'s
|
||
`qwen-classifier` comment (unfused attention buffers ballooned unexpectedly at 65536 ctx; flash-attn
|
||
fixed it). `--flash-attn on` is already the deployed default for the main model, so the linear-ish
|
||
regime applies, but no primary source gives an exact formula for this buffer size at 1M context — the
|
||
+3-6 GiB band is this doc's estimate based on the ratio observed in that in-repo incident, not a
|
||
llama.cpp-documented number. Budget for the high end of that range when sizing hardware.
|
||
|
||
**Bottom line:** at `q4_0` KV (the most aggressive, already-proven-in-this-repo tier), 500k×2 /
|
||
1M needs **~36-39 GB** total VRAM for the 27B model alone. That does not fit one 32GB card at any
|
||
precision — it needs at least two 32GB-class cards, or one ≥40GB card. At `q8_0` (the precision this
|
||
repo already runs in production for quality reasons), budget **~51-54 GB** — two 32GB cards (64GB
|
||
pooled) clears this with room to spare; a single 48GB-class card would not.
|
||
|
||
(§11 below extends this table to higher weight-quant tiers — Q6_K_XL, Q8_0, BF16 — for users who want
|
||
better output quality than `Q4_K_XL`, and to a Flash-Next alternative architecture; see §11.6-§11.7.)
|
||
|
||
---
|
||
|
||
## 5. CPU/motherboard: which Threadripper generations give PCIe 4.0, and how many lanes for GPUs
|
||
|
||
AMD's own product/chipset pages, cross-checked against the launch reviews that quote them directly:
|
||
|
||
| Platform | Socket | PCIe generation | Total CPU-provided lanes |
|
||
|---|---|---|---|
|
||
| Threadripper 1000/2000 series ("1920X", "2950X", etc.) | **TR4** | **PCIe 3.0 only** | 60-64 |
|
||
| Threadripper 3000 series (3960X/3970X/3990X) | **sTRX4** | **PCIe 4.0** | 64 |
|
||
| Threadripper 7000 series (non-PRO) | sTR5 | PCIe 5.0 (48 lanes) + PCIe 4.0 (24-32 lanes) | ~72-80 |
|
||
| Threadripper PRO 3000WX/5000WX | sWRX8 | PCIe 4.0 | **128** |
|
||
| Threadripper PRO 7000WX | sTR5 (WRX90) | PCIe 5.0 (128 lanes) + a few PCIe 3.0 | **128** |
|
||
|
||
Sources: [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html)
|
||
(sWRX8/socket listing), corroborated by
|
||
[Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch coverage](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2)
|
||
(*"the 3rd Gen TR CPUs carry the same 64 PCIe lanes but double bandwidth by moving from Gen 3.0 to
|
||
Gen 4.0"* — explicit confirmation TR4/1000-2000-series is PCIe 3.0 while sTRX4/3000-series is PCIe
|
||
4.0), [PCWorld — Threadripper PRO launch](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html)
|
||
(*"128 PCIe lanes"* for PRO).
|
||
|
||
**This directly matters for the user's plan.** "An older Threadripper... for more PCIe lanes" is
|
||
ambiguous between two real, very different chips:
|
||
|
||
- **TR4 (1000/2000 series)** — cheapest used option, but **PCIe 3.0** — does not meet the user's own
|
||
stated PCIe 4.0 requirement, and PCIe 3.0 x8 per GPU roughly halves inter-GPU/host transfer
|
||
bandwidth (matters more for training/tensor-parallel than for llama.cpp's inference-time layer
|
||
splitting, but still a real downgrade vs. the R9700's native PCIe 5.0).
|
||
- **sTRX4 (3000 series, non-PRO)** — the correct "older Threadripper with PCIe 4.0" target: 64 lanes,
|
||
4-5 years old, real used-market availability, no PRO price premium.
|
||
- **Threadripper PRO (3000WX/5000WX)** doubles the lane count to 128 but at meaningfully higher used
|
||
cost (workstation-tier, lower volume, sWRX8 boards are pricier than sTRX4/TRX40 boards) — worth it
|
||
only if 6 full-bandwidth (x16) GPU slots are actually needed; at x8-per-card (adequate for inference)
|
||
64 lanes already covers 6 GPUs with lanes to spare for NVMe/chipset.
|
||
|
||
**Lane budget for 6 GPUs on sTRX4 (64 lanes), estimated (no vendor spec gives a topology this
|
||
specific — treat this bullet as an estimate):** typical sTRX4 boards reserve ~4 lanes for the
|
||
chipset uplink and commonly wire 1-2 M.2 slots directly to the CPU (4 lanes each) — so realistic
|
||
GPU-available lanes land around 44-52 of the 64, i.e. **6 GPUs at x8 electrical each (48 lanes) is
|
||
plausible but board-model-dependent**; x16-each for 6 cards is not possible on 64 lanes regardless of
|
||
board. x8 electrical is not a meaningful inference-speed penalty for llama.cpp (weights are loaded
|
||
once; the ongoing per-token traffic across PCIe is small compared to compute), so this is an
|
||
acceptable tradeoff, not a real bottleneck for this workload.
|
||
|
||
---
|
||
|
||
## 6. Power budget vs. the ~200W / one spare 8-pin headroom
|
||
|
||
Official/vendor TDPs:
|
||
|
||
| Card | TDP | Source |
|
||
|---|---|---|
|
||
| GTX 1080 (owned) | 180W, one 8-pin | [confirmed 180W, PCIe 3.0 x16, 1× 8-pin](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu) — spec matches NVIDIA's own launch figures reported across multiple outlets incl. Tom's Hardware |
|
||
| RTX 2080 (owned) | 215W | Cross-checked across gpuzoo/cputronic/notebookcheck spec pages, consistent at 215W |
|
||
| GT 710 (owned) | ~19W, **no external power connector** (slot power only) | [MSI/EVGA/Zotac GT 710 spec pages](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification) |
|
||
| RTX 3060 12GB (candidate purchase) | 170W, one 8-pin | [NVIDIA-confirmed 170W TDP, one 8-pin connector](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/) |
|
||
| RTX 3090 24GB (candidate purchase) | 350W, two 8-pin, [NVIDIA's own RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/) lists 350W and a 750W PSU minimum | NVIDIA official |
|
||
| R9700 32GB (already deployed / "more of the same") | 300W (per this repo's `server-planing.md`, consistent with AMD's own R9700 product page framing it as a 300W-class card) | in-repo prior research |
|
||
|
||
**Against the stated ~200W / one spare 8-pin budget:**
|
||
|
||
- Adding **one RTX 3060 12GB** (170W, one 8-pin) is the *only* candidate in this list that fits the
|
||
stated headroom as-is — it uses the one spare connector and stays under 200W.
|
||
- Adding the already-owned **GTX 1080** (180W) as the classifier's dedicated card also just barely
|
||
fits (180W ≤ 200W, one 8-pin) — this is a genuinely free option since the card is already owned and
|
||
its power draw is within budget, unlike every purchase candidate below.
|
||
- Adding the already-owned **RTX 2080** (215W) **exceeds** the stated 200W headroom by 15W — technically
|
||
over budget on paper, though real-world draw is usually a bit under rated TDP; flag it as marginal,
|
||
not safely fitting.
|
||
- Adding a **second R9700** (300W) or an **RTX 3090** (350W, needs two 8-pin — the user has only one
|
||
spare) both blow well past the current power budget on both watts and connector count.
|
||
- **The GT 710 draws no meaningful power (~19W, no PCIe power connector at all)** — it is free from a
|
||
power-budget standpoint regardless of what else is added.
|
||
|
||
**PSU upgrade trigger:** the very first stage that adds *any* GPU beyond a GTX 1080-class card (180W,
|
||
one 8-pin) or an RTX 3060 12GB (170W, one 8-pin) exhausts the stated headroom. Any stage that reaches
|
||
for a second 32GB-class card (R9700 or equivalent) or any 300W+ card **requires a PSU upgrade before
|
||
that stage**, not after — see the roadmap table in §7 for exactly which stage that is.
|
||
|
||
---
|
||
|
||
## 7. Mixed-GPU feasibility: ROCm + CUDA, and is the GT 710 usable at all
|
||
|
||
**ROCm and CUDA are different llama.cpp builds, but that's exactly the pattern already in this
|
||
repo.** `ghcr.io/ggml-org/llama.cpp` publishes both `server-rocm` and `server-cuda` as separate,
|
||
independently-built image tags (confirmed present on the [ggml-org container registry](https://github.com/orgs/ggml-org/packages/container/llama.cpp)
|
||
and documented in [`docs/docker.md`](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md)
|
||
— *"server-cuda: Same as `server` but compiled with CUDA support"*, *"server-rocm: Same as `server`
|
||
but compiled with ROCm support"*). You cannot mix backends inside one process/container, but you
|
||
**can** run one `server-rocm` container pinned to the R9700 and a separate `server-cuda` container
|
||
pinned to an NVIDIA card, simultaneously, on the same host — this is architecturally identical to
|
||
today's `llama-server` + `qwen-classifier` two-container split in `docker-compose.yml`, just with a
|
||
different image tag for the NVIDIA-backed service and NVIDIA's container runtime (`nvidia-container-toolkit`
|
||
+ `--gpus` / device reservation, the CUDA-world equivalent of this repo's `/dev/kfd`+`/dev/dri`+
|
||
numeric-GID ROCm pattern documented in
|
||
[`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md)). None of that doc's ROCm-specific
|
||
findings (the `GPU_MAX_HW_QUEUES=1` MES firmware workaround, the numeric-GID `group_add` fix) apply to
|
||
an NVIDIA/CUDA container — those are ROCm-stack-specific bugs, not general multi-GPU-container issues.
|
||
|
||
**Is this an implicit AMD→NVIDIA rebuild, or additive?** Worth surfacing explicitly since the two
|
||
source plans conflict on this: `server-planing.md` is an AMD-only, ROCm-only 4×R9700 plan. This
|
||
ticket's plan is **additive/mixed** — keep the R9700 running the main model under ROCm, and bolt on
|
||
NVIDIA cards under CUDA for secondary roles (classifier, or a second inference GPU for the big model
|
||
if going the "more of the same type" route means buying NVIDIA instead of more R9700s). Both are
|
||
internally consistent, but they are different end-states — flag this choice back to the user rather
|
||
than assuming one.
|
||
|
||
**Splitting the *main* 27B model itself across mixed AMD+NVIDIA silicon in one process is not
|
||
possible** — llama.cpp's multi-GPU tensor-split only works within a single backend build. To use
|
||
both an R9700 and an NVIDIA card for the *same* model's layers, all the compute-hosting cards need to
|
||
be the same backend (all-ROCm or all-CUDA) in that one process. This is why §5's roadmap treats "add
|
||
GPU capacity to the main model" and "add a GPU for the classifier" as separable purchases with
|
||
different backend constraints, not a single mixed pool.
|
||
|
||
**Is the GT 710 usable for anything in this pipeline? No.** Reasoning:
|
||
|
||
- 1GB VRAM cannot hold any meaningful fraction of either model's weights (17.6 GB / 2.4-4.3 GB) —
|
||
even a handful of transformer layers at Q4 quantization exceeds 1GB.
|
||
- It's Kepler-generation silicon (192 CUDA cores, no tensor cores) — llama.cpp's CUDA backend
|
||
technically supports pre-Turing cards, but at this VRAM size there's nothing to usefully offload.
|
||
- It draws power from the PCIe slot only, no external connector — genuinely free to keep installed.
|
||
- **Plausible actual use: dedicate it as the box's display-output card**, so every compute-capable
|
||
GPU (R9700, and whichever NVIDIA cards get added) can be fully headless/compute-only with none of
|
||
their VRAM or a display output tied up driving a monitor — a real, if minor, use for it. This is
|
||
this doc's own inference from the spec facts above, not a claim found in any primary source.
|
||
|
||
---
|
||
|
||
## 8. GPU market pricing vs. the $20-30/GB assumption
|
||
|
||
| Card | VRAM | Backend | Current used-market price (estimate — see caveat) | $/GB |
|
||
|---|---|---|---|---|
|
||
| RTX 3060 12GB | 12GB | CUDA | ~$240-300 used (eBay listings, [gpupoet.com tracker](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060): *"from $239"*, [eBay live listings](https://www.ebay.com/shop/rtx-3060-12gb) averaging ~$488 asking but with a $239 floor) | **~$20-25/GB** — matches the stated assumption |
|
||
| RTX 3090 24GB | 24GB | CUDA | ~$1,010-1,050 used ([bestvaluegpu.com Sep 2026 tracker](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/), [xda-developers coverage](https://www.xda-developers.com/used-rtx-3090-still-best-for-local-ai-in-value/)) | **~$42-44/GB** — well above the stated assumption |
|
||
| R9700 32GB ("more of the same type") | 32GB | ROCm | **New only — $1,299 MSRP**, street price $1,400-1,585 as of this research ([overclock3d](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/), [pricehistory.app tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm)) — too recent a release (2026) for a used market to exist yet | **~$41-50/GB, and not a used-market price at all** |
|
||
|
||
**Caveat on all three price figures:** these are live marketplace asking-price snapshots pulled via
|
||
web search on 2026-09-11, not sold-price data or a vendor spec sheet — treat as directional, not
|
||
exact. eBay asking prices in particular run above realized sale prices.
|
||
|
||
**Correction to the user's stated assumption:** $20-30/GB is a good estimate specifically for
|
||
**last-generation mainstream used cards** (RTX 3060 12GB fits it almost exactly) but **not** for
|
||
high-VRAM flagship cards like the RTX 3090 (~1.5-2× that rate) or for "more of the same type" R9700
|
||
units, which aren't used-market at all yet and sit even higher per GB than the 3090. If the plan is
|
||
"cheapest path to more VRAM," multiple RTX 3060 12GB cards (or similar mid-tier used cards) beat one
|
||
RTX 3090 on $/GB, at the cost of needing more PCIe slots and more total wattage/connectors to reach
|
||
the same aggregate VRAM — which is exactly the tradeoff the Threadripper lane-count plan in §5 is
|
||
for.
|
||
|
||
---
|
||
|
||
## 9. Step-by-step roadmap
|
||
|
||
All "resulting max context" figures assume `--parallel 2` and the KV precision stated; per §2, the
|
||
`--ctx-size` value shown is the *total* (pre-division) value to pass to llama-server.
|
||
|
||
| Stage | Hardware change | Est. cost | Backend | Usable VRAM (main-model pool) | Max context @ parallel=2 (`q4_0` KV) | PSU upgrade triggered? |
|
||
|---|---|---|---|---|---|---|
|
||
| **0 (current)** | 1× R9700 32GB, in production | $0 | ROCm | 32GB (shared with classifier) | ~131K/slot today at `q8_0` KV (262144 total ÷ 2) | No |
|
||
| **1 — classifier isolation** | Move classifier onto the already-owned **GTX 1080** (180W, own container, `server-cuda`), freeing the R9700 entirely for the main model. Matches the existing dual-model pattern qwen-code's own docs describe (§10) and this repo's `qwen-classifier` service already implements, just on separate silicon instead of a shared card. | $0 (already owned) | ROCm (main) + CUDA (classifier) | R9700's full 32GB now available to the main model alone | ~262K/slot @ `q8_0` (unchanged ctx-size, no more classifier contention) | **No** — 180W GTX 1080 fits the stated ~200W/one-8-pin headroom |
|
||
| **2 — second big-model GPU** | Add **one more 32GB-class card** for the main model. Cheapest correct-backend option: a second R9700 (~$1,300-1,585 new, ROCm, same backend as the first — required if tensor-splitting one model across two cards) | ~$1,300-1,585 | ROCm | 64GB pooled | `--ctx-size 500000 --parallel 1` fits at `q4_0` (~33GB) or `q8_0` (~48GB, tight but fits in 64GB) — **not yet 500k×2** | **Yes** — 300W card, no spare 8-pin left after stage 1 |
|
||
| **3 — reach 500k × 2 / 1M stretch** | No further hardware if stage 2's 64GB pool is used with `--cache-type-k/v q4_0`: `--ctx-size 1000000 --parallel 2` needs ~33-39GB (§4), fits inside 64GB with real headroom for the compute buffer. If `q8_0` KV is required instead (this repo's current quality bar for the main model), the ~51-54GB need is tight-to-marginal on 64GB — a **third** 32GB card (~96GB pool) removes the risk. | $0 (reuses stage 2) or +$1,300-1,585 for a 3rd card if `q8_0` KV is required | ROCm | 64GB (q4_0 case) or 96GB (q8_0 case) | **500k×2 parallel achieved**, and the 1M stretch goal is the *same config* with `--parallel 1` instead of 2 (§2) | Already upgraded at stage 2 |
|
||
| **4 — optional CPU/lane platform swap** | Only needed if the plan is to keep scaling past 2-3 big cards, or to add several small used cards (RTX 3060 12GB) for extra headroom/throughput rather than raw ctx-size. Swap to **non-PRO Threadripper 3000-series (sTRX4)** — 64 PCIe 4.0 lanes, ~x8-per-slot for up to 6 GPUs (§5). Threadripper PRO 3000WX/5000WX (128 lanes) only if x16-per-card matters or 6+ full-bandwidth slots are wanted. | Used sTRX4 CPU+board: roughly $400-800 combined on the used market (not independently priced in this pass — **estimate**, not cited) | n/a (platform only) | n/a | n/a | Independent of GPU wattage — driven by whatever GPU count/wattage stage 5+ adds |
|
||
| **5+ — scale-out via small used cards** | Add RTX 3060 12GB units (~$20-25/GB, the assumption that actually holds, §8) instead of more 32GB flagship cards, once lane count (stage 4) supports it — useful for extra parallel slots / throughput beyond the 500k×2 target rather than for raising ctx-size further (500k×2/1M is already met at stage 3). | ~$240-300/card | CUDA (separate container per §6) | +12GB pooled per card, but on a *different backend* from the ROCm main model — usable for extra classifier/small-model capacity or a separate CUDA-backend llama-server instance, not as additional tensor-split VRAM for the ROCm main model | Unchanged for the main model; adds parallel capacity elsewhere | Yes, cumulative — each additional 170W card needs PSU headroom stage 2 already consumed |
|
||
|
||
**Where the existing dual-model pattern sits in this roadmap:** it's stage 1, and it's free. The
|
||
qwen-code docs pattern (main model + a small, always-resident, non-thinking fast/classifier model —
|
||
see §10) is already implemented in this repo; the only roadmap-relevant change is *which GPU* the
|
||
classifier sits on, moving it off the R9700 entirely onto an already-owned NVIDIA card frees the
|
||
R9700's full 32GB for the 500k×2/1M push instead of splitting it with the classifier as happens
|
||
today.
|
||
|
||
### 9.1 Upgrade path, as diagrams
|
||
|
||
Diagram form of the same §9 table and §11.9's dense-vs-Flash-Next call — nothing new is claimed here,
|
||
this is a visual index back into the cited sections above.
|
||
|
||
**Stage-by-stage hardware path** (PSU-upgrade triggers and target reached called out inline):
|
||
|
||
```mermaid
|
||
flowchart TD
|
||
S0["Stage 0 — today<br/>1x R9700 32GB, ROCm<br/>classifier shares the card<br/>$0"]
|
||
S1["Stage 1 — classifier isolation<br/>+ GTX 1080 (owned, 180W, CUDA)<br/>R9700 freed for main model<br/>$0 · PSU OK (180W fits ~200W headroom)"]
|
||
S2["Stage 2 — 2nd big-model GPU<br/>+1x R9700 32GB (ROCm)<br/>64GB pooled<br/>~$1,300-1,585 · PSU UPGRADE REQUIRED (300W, no 8-pin left)"]
|
||
S3q4["Stage 3a — q4_0 KV<br/>--ctx-size 1,000,000 --parallel 2<br/>~33-39GB, fits in 64GB<br/>$0 (reuses stage 2)"]
|
||
S3q8["Stage 3b — q8_0 KV (current prod quality)<br/>~51-54GB, tight on 64GB<br/>+1x R9700 -> 96GB removes risk<br/>+~$1,300-1,585"]
|
||
TARGET(["500k x2 parallel reached<br/>= 1M stretch goal, same VRAM<br/>(--parallel 1 vs 2 is a config flag, §2)"])
|
||
S4["Stage 4 — platform swap (optional)<br/>sTRX4 Threadripper 3000, 64 PCIe4 lanes<br/>only needed past 2-3 big cards<br/>~$400-800 (estimate, §9)"]
|
||
S5["Stage 5+ — scale out<br/>+RTX 3060 12GB cards (CUDA, separate backend)<br/>extra parallel/throughput, not more ctx-size<br/>~$240-300/card · PSU upgrade each card"]
|
||
|
||
S0 --> S1 --> S2
|
||
S2 --> S3q4 --> TARGET
|
||
S2 --> S3q8 --> TARGET
|
||
TARGET -.->|"only if scaling past this"| S4 --> S5
|
||
|
||
style TARGET fill:#2e7d32,color:#fff,stroke:#1b5e20
|
||
style S2 fill:#8a5a00,color:#fff,stroke:#5c3d00
|
||
style S5 fill:#8a5a00,color:#fff,stroke:#5c3d00
|
||
```
|
||
|
||
**Model choice, and the one open question that could change it** (§11.9):
|
||
|
||
```mermaid
|
||
flowchart TD
|
||
Q{"Goal: 500k-1M ctx<br/>within a $20-30/GB VRAM budget?"}
|
||
D["Dense Qwen3.8-27B<br/>17.6-54.7GB weights (Q4_K_XL-BF16)<br/>reaches 500k x2 on 2-3 cards,<br/>500k x4 on 2-6 cards depending on quant<br/>(§11.7 tables)"]
|
||
F{"Try --n-cpu-moe:<br/>offload MoE experts to system RAM?<br/>(untested for this model, §11.8)"}
|
||
FBAD["Flash-Next, all-GPU weights<br/>111-354GB just for weights<br/>needs 4-13 cards before any KV cost<br/>NOT recommended at this budget (§11.9)"]
|
||
FGOOD["Flash-Next, experts in system RAM<br/>GPU VRAM could shrink a lot<br/>2.67x cheaper KV/token becomes relevant<br/>UNVERIFIED — prototype on real server first"]
|
||
CAVEAT["+ real caveat either way:<br/>PR #27742 flags unverified conv branch,<br/>3% QSA divergence, prefill-pos-0-only PLE<br/>(§11.1) — dense model carries no such flag"]
|
||
|
||
Q --> D
|
||
Q -->|"considering Flash-Next instead"| F
|
||
F -->|"works well"| FGOOD
|
||
F -->|"doesn't help / untested"| FBAD
|
||
FGOOD --> CAVEAT
|
||
FBAD --> CAVEAT
|
||
|
||
style D fill:#2e7d32,color:#fff,stroke:#1b5e20
|
||
style FBAD fill:#8a1c1c,color:#fff,stroke:#5c1212
|
||
style FGOOD fill:#8a5a00,color:#fff,stroke:#5c3d00
|
||
```
|
||
|
||
---
|
||
|
||
## 10. Qwen-code's own docs on the fast-model/classifier pattern
|
||
|
||
Fetched directly per the user's link:
|
||
[qwenlm.github.io/qwen-code-docs/en/users/overview/](https://qwenlm.github.io/qwen-code-docs/en/users/overview/)
|
||
— **the overview page itself does not describe the dual-model/classifier pattern**; it only covers
|
||
single-model-provider setup (Alibaba ModelStudio / third-party / custom provider), one model at a
|
||
time. The actual fast-model/classifier documentation lives on the **Auto Mode** page instead, which
|
||
this repo's own `fast-model-choice.md` already fetched and cited in detail:
|
||
[qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/) —
|
||
summary (see `fast-model-choice.md` §1 for the full quote): a two-stage classifier gate, both stages
|
||
using "your configured fast model (`/model --fast`)", Stage 1 a ~300ms `{shouldBlock}`-only check,
|
||
Stage 2 a ~3-5s chain-of-thought review that only runs on a Stage-1 block. Nothing in either page
|
||
gives a recommended *context size* or *model size* for the fast model beyond what's implied by that
|
||
latency budget — this repo's own prior research (`fast-model-choice.md`) derived the actual context
|
||
requirement from qwen-code's source code instead (`packages/core/src/permissions/classifier-transcript.ts`),
|
||
since the docs pages don't state one. No new information changes that prior doc's conclusion; this
|
||
section exists to confirm the overview page was checked directly as instructed and doesn't contradict
|
||
or add to it.
|
||
|
||
---
|
||
|
||
## 11. Alternative: Qwen3.8-Flash-Next (MoE, hybrid attention)
|
||
|
||
The user also wants to weigh switching (or adding) **Qwen3.8-Flash-Next** — a 125B-total/6B-active MoE
|
||
with a hybrid recurrent-attention architecture — against staying on dense Qwen3.8-27B, and separately
|
||
wants this section to cover **going up in weight quant** (Q4_K_XL → Q6_K_XL → Q8_0 → BF16/fp16) for
|
||
*both* models, not just Q4. Feasibility first, since it gates everything else.
|
||
|
||
### 11.1 Feasibility verdict: supported, but immature — read before trusting any number below
|
||
|
||
Checked directly against the primary sources the task named:
|
||
|
||
- **llama.cpp mainline support exists.** [PR #27742](https://github.com/ggml-org/llama.cpp/pull/27742)
|
||
("model: add Qwen3.8-Flash-Next (qwen4exp)") was **merged into `master` on 2026-08-27** by ngxson.
|
||
It adds the full architecture: Gated DeltaNet layers (sigmoid-gated linear attention), QSA
|
||
("Qwen Sparse Attention", operating at micro-block granularity), hyper-connections, and the PLE
|
||
n-gram embedding table. `llama.cpp`'s own docs list CPU/CUDA/Metal/ROCm as supported backends for
|
||
it — this is not a CUDA-only feature.
|
||
- **This repo's pinned image is a floating tag, not a version pin.** `docker-compose.yml` runs
|
||
`ghcr.io/ggml-org/llama.cpp:server-rocm` with no date/digest suffix — a rolling "latest ROCm server
|
||
build" tag, not a release version. The merge is from 2026-08-27, and today is 2026-09-11 (~2 weeks
|
||
later), so a **fresh pull** of `server-rocm` should include it — but whatever image is already
|
||
cached/running on the R9700 box may predate the merge. **Action before touching this model on the
|
||
server: `docker compose pull llama-server` and check the startup log's build/commit banner is dated
|
||
on/after 2026-08-27**, not just "the tag says server-rocm."
|
||
- **Real, primary-source-flagged immaturity — this is the part that should temper enthusiasm.** The
|
||
PR's own description/review discussion states: *"The conv branch itself is still numerically
|
||
unverified because the fixture zeroes its weights"*; QSA sparse attention *"diverges on 3 percent of
|
||
positions"* above its budget threshold; the PLE depthwise convolution *"is exact only for a prefill
|
||
that starts at position 0"* (i.e. correctness is not guaranteed once `--cache-reuse`/prompt-caching
|
||
is in play — a flag this repo already turns on for the dense model per the latest commit). None of
|
||
that is disqualifying, but it is a primary-source admission that this is a fresh, not-fully-verified
|
||
implementation, not a mature, widely-battle-tested one like the dense Qwen3.8-27B path.
|
||
- **Multi-slot serving needs an explicit new flag.** The same PR states: *"`set_input_qsa` asserted
|
||
`n_stream == 1`, so llama-server could not serve this model with more than one slot unless `-kvu`
|
||
was passed."* Per the server README (§2), `--kv-unified`/`-kvu` defaults to enabled **only when slot
|
||
count is auto** (`-1`). This repo's compose file sets `--parallel ${LLAMA_PARALLEL:-2}` **explicitly**
|
||
(not auto) — so adopting Flash-Next with `--parallel` > 1 requires **adding `--kv-unified` (or
|
||
`-kvu`) to the launch flags**, a real deploy-time change, not something that "just works" by copying
|
||
today's flag set onto a new model file.
|
||
|
||
**Verdict: yes, runnable** on this repo's backend (ROCm, mainline, no dev branch needed) as long as the
|
||
image is pulled after 2026-08-27 and `-kvu` is added for multi-slot use — but treat it as
|
||
**usable-with-caution**, not a drop-in swap, given the PR author's own unresolved-correctness notes.
|
||
|
||
### 11.2 Architecture, verified against `config.json` directly
|
||
|
||
Fetched from `Qwen/Qwen3.8-Flash-Next`'s `config.json` (unsloth's GGUF repo repackages the same base
|
||
model): `num_hidden_layers=48`, `hidden_size=2560`, `num_attention_heads=24`, `num_key_value_heads=2`,
|
||
`head_dim=256`, `max_position_embeddings=262144` (same native/extensible-to-1M framing as the dense
|
||
model — same YaRN quality caveat from §3 applies here too, unverified past native length), `num_experts=512`,
|
||
`num_experts_per_tok=10`, and the linear-attention head config: `linear_num_key_heads=16`,
|
||
`linear_num_value_heads=48`, `linear_key_head_dim=128`, `linear_value_head_dim=128`.
|
||
|
||
Layer pattern (confirmed both from the model card's own description and `config.json`'s
|
||
`full_attention_interval=4`): every 4th layer is full/QSA attention, the other 3 are Gated DeltaNet —
|
||
**12 of 48 layers grow a real KV cache; the other 36 have a fixed-size recurrent state that does not
|
||
grow with context length.** (24% full-attention layers vs. the dense model's 16-of-64 = 25% — similar
|
||
ratio, but the *absolute* per-layer KV cost differs because `num_key_value_heads` is 2 here vs. 4 on
|
||
the dense model — see below.)
|
||
|
||
### 11.3 Per-token growing-KV-cache cost
|
||
|
||
```
|
||
12 full-attention layers × 2 (K+V) × 2 kv_heads × 256 head_dim × 2 bytes (fp16) = 24 KiB/token
|
||
```
|
||
|
||
| Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 |
|
||
|---|---|---|---|
|
||
| 262,144 (native) | ~6.0 GiB | ~3.0 GiB | ~1.5 GiB |
|
||
| 500,000 | ~11.4 GiB | ~5.7 GiB | ~2.9 GiB |
|
||
| **1,000,000 (500k×2, or 1M stretch)** | **~22.9 GiB** | **~11.4 GiB** | **~5.7 GiB** |
|
||
| **2,000,000 (500k×4)** | **~45.8 GiB** | **~22.9 GiB** | **~11.4 GiB** |
|
||
|
||
`--cache-type-k/v` are the same generic llama.cpp KV-cache-quantization flags used elsewhere in this
|
||
doc; nothing in the PR or the server README suggests they're handled differently for the 12
|
||
full-attention layers of a hybrid model — they quantize the same growing K/V buffers as on a plain
|
||
transformer. (No primary source explicitly confirms this for *this* architecture specifically — flagged
|
||
as a reasonable extrapolation, not a directly-cited fact, same caveat class as this doc's other
|
||
estimates.)
|
||
|
||
### 11.4 Fixed (non-growing) recurrent state — Gated DeltaNet layers
|
||
|
||
The 36 Gated DeltaNet layers each keep a fixed-size recurrent state (an outer-product-style
|
||
key×value matrix per head) that does **not** scale with context length — only with slot/sequence
|
||
count. Sized from `config.json`'s linear-attention head params:
|
||
|
||
```
|
||
36 layers × linear_num_value_heads(48) × linear_key_head_dim(128) × linear_value_head_dim(128) × 4 bytes (fp32 state)
|
||
≈ 36 × 48 × 128 × 128 × 4 bytes ≈ 108 MiB per slot
|
||
```
|
||
|
||
This is **this doc's own derivation from the published head-dimension params, not a value pulled
|
||
directly from llama.cpp source or docs** — the PR text confirms the state exists per-stream/per-slot
|
||
but doesn't publish an exact byte formula, so treat the ~108 MiB/slot figure as an estimate, medium
|
||
confidence. Even at 4 parallel slots that's under half a gigabyte — **negligible** next to both the
|
||
growing KV cache (GBs) and the weights (tens to hundreds of GB) computed below. The headline
|
||
implication holds regardless of the exact multiplier: Flash-Next's "big memory line item" is the MoE
|
||
weights, not the attention state of any kind.
|
||
|
||
### 11.5 Magnitude vs. the dense model — how much cheaper is KV, really
|
||
|
||
At the same total ctx-size, Flash-Next's growing KV cache is **24 KiB/token vs. the dense model's
|
||
64 KiB/token — 2.67× smaller**, i.e. Flash-Next's KV budget is **37.5%** of the dense model's at
|
||
identical context length. This is a real, significant win *for the KV-cache line item specifically* —
|
||
but see §11.9: it's a much smaller slice of a much bigger total, because the weights move the other
|
||
way by a far larger factor.
|
||
|
||
### 11.6 Weight sizes — verified from each unsloth GGUF repo's actual file listing
|
||
|
||
Fetched directly from the HF file trees (not estimated from ratios), current as of this research pass:
|
||
|
||
| Quant tier | Qwen3.8-27B (dense) | Qwen3.8-Flash-Next (MoE) |
|
||
|---|---|---|
|
||
| Q4_K_XL (`UD-Q4_K_XL`) | **17.6 GB** (existing baseline) | **111.4 GB** (4 parts: 10.9MB + 49.9GB + 49.4GB + 12.1GB) |
|
||
| Q6_K_XL (`UD-Q6_K_XL`) | **25.3 GB** | **169 GB** (6 parts) |
|
||
| Q8_0 | **29 GB** | **188 GB** (6 parts) |
|
||
| BF16/fp16 | **54.67 GB** (50GB + 4.67GB, 2 parts) | **354 GB** (8 parts) |
|
||
|
||
Sources: [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main)
|
||
and its `BF16/` subfolder; [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main)
|
||
and its `UD-Q4_K_XL/`, `UD-Q6_K_XL/`, `Q8_0/`, `BF16/` subfolders (per-file sizes summed). The
|
||
preliminary Q8_0 figure floated before this research pass (~192GB) was slightly high — the real
|
||
listing sums to **188 GB**; everything else in the preliminary list was accurate to within rounding.
|
||
|
||
**The weight-quant axis and the KV-cache-quant axis are independent knobs.** Raising weight quality
|
||
(Q4_K_XL → BF16) does not require raising `--cache-type-k/v` — the two flags are unrelated, and this
|
||
repo already proves that pattern works (`q8_0` KV cache is deployed today against `Q4_K_XL` weights).
|
||
A user chasing **maximum output quality** can run e.g. **BF16 weights + `q4_0` KV cache** — full-precision
|
||
weights for quality, still-compressed KV for context budget — or any other combination in the tables
|
||
below; nothing about picking a higher weight quant forces a matching KV precision.
|
||
|
||
### 11.7 Total VRAM: does it fit, across quant tiers and both parallelism targets
|
||
|
||
All totals = weights + growing KV cache + an estimated **+3-6 GB** compute-buffer/runtime overhead
|
||
(same estimate band as §4, carried over — not re-derived for this architecture; flagged medium
|
||
confidence there too). "Cards" = ceil(total ÷ 32GB), i.e. how many R9700-class 32GB cards it takes.
|
||
|
||
#### Dense Qwen3.8-27B — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`)
|
||
|
||
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|
||
|---|---|---|---|
|
||
| Q4_K_XL (17.6GB) | ~79-85GB (**3**) | ~48-54GB (**2**) | ~33-39GB (**2**) |
|
||
| Q6_K_XL (25.3GB) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) | ~44-47GB (**2**) |
|
||
| Q8_0 (29GB) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) | ~47-50GB (**2**) |
|
||
| BF16 (54.67GB) | ~119-122GB (**4**) | ~88-91GB (**3**) | ~73-76GB (**3**) |
|
||
|
||
#### Dense Qwen3.8-27B — 500k × 4 parallel (`--ctx-size 2,000,000`)
|
||
|
||
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|
||
|---|---|---|---|
|
||
| Q4_K_XL (17.6GB) | ~143-146GB (**5**) | ~82-85GB (**3**) | ~51-54GB (**2**) |
|
||
| Q6_K_XL (25.3GB) | ~150-153GB (**5**) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) |
|
||
| Q8_0 (29GB) | ~154-157GB (**5**) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) |
|
||
| BF16 (54.67GB) | ~180-183GB (**6**) | ~119-122GB (**4**) | ~88-91GB (**3**) |
|
||
|
||
#### Qwen3.8-Flash-Next — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`)
|
||
|
||
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|
||
|---|---|---|---|
|
||
| UD-Q4_K_XL (111.4GB) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) | ~120-123GB (**4**) |
|
||
| UD-Q6_K_XL (169GB) | ~195-198GB (**7**) | ~183-186GB (**6**) | ~178-181GB (**6**) |
|
||
| Q8_0 (188GB) | ~214-217GB (**7**) | ~202-205GB (**7**) | ~197-200GB (**7**) |
|
||
| BF16 (354GB) | ~357-360GB (**12**) | ~357-360GB (**12**) | ~357-360GB (**12**) |
|
||
|
||
#### Qwen3.8-Flash-Next — 500k × 4 parallel (`--ctx-size 2,000,000`)
|
||
|
||
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|
||
|---|---|---|---|
|
||
| UD-Q4_K_XL (111.4GB) | ~160-163GB (**6**) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) |
|
||
| UD-Q6_K_XL (169GB) | ~218-221GB (**7**) | ~195-198GB (**7**) | ~183-186GB (**6**) |
|
||
| Q8_0 (188GB) | ~234-237GB (**8**) | ~211-214GB (**7**) | ~199-202GB (**7**) |
|
||
| BF16 (354GB) | ~397-400GB (**13**) | ~377-380GB (**12**) | ~366-369GB (**12**) |
|
||
|
||
(Flash-Next's KV precision barely moves the total at any weight quant above `UD-Q6_K_XL` — the weights
|
||
so dominate the budget that KV quantization stops mattering for the "how many cards" question. This
|
||
is the clearest signal in this whole section: for Flash-Next, the weight-quant choice is the entire
|
||
hardware-sizing decision; for the dense model, KV precision still matters a lot.)
|
||
|
||
### 11.8 CPU MoE-expert offload — the one lever that could change this calculus
|
||
|
||
Flash-Next is a 512-expert/10-active-per-token MoE, and llama.cpp has a purpose-built flag for exactly
|
||
this shape of model, confirmed directly from the server README: **`--n-cpu-moe`** — *"keep the Mixture
|
||
of Experts (MoE) weights of the first N layers in the CPU"* — plus the more general
|
||
**`--override-tensor`** (*"override tensor buffer type"*, pattern-matched by tensor name) that the same
|
||
flag is built on top of. Both are generic, architecture-agnostic llama.cpp mechanisms (they match on
|
||
tensor name patterns, not model type), so there's no reason to expect them not to apply to Flash-Next's
|
||
MoE tensors specifically — but this pass found **no primary source that has actually tested
|
||
`--n-cpu-moe` against this specific qwen4exp architecture**, so treat "it works here" as plausible,
|
||
not confirmed.
|
||
|
||
If it does work as expected, this changes the whole weight-VRAM picture in §11.7: the ~90-95% of
|
||
Flash-Next's weight footprint that's MoE expert tensors could live in system RAM while attention
|
||
projections, the shared/non-expert tensors, and the full KV cache stay on GPU — meaning a much smaller
|
||
GPU-VRAM number than the "all weights on GPU" tables above, at the cost of PCIe/RAM-bandwidth-bound
|
||
inference speed for whichever experts get selected per token (this repo has no benchmark of that
|
||
tradeoff, and it's highly system-RAM-bandwidth-dependent, so no number is given here — flagged as an
|
||
escape hatch worth prototyping directly on the server, not something this research values responsibly
|
||
without a real test run).
|
||
|
||
### 11.9 Net recommendation: dense Qwen3.8-27B vs. Flash-Next, for this user's stated goal
|
||
|
||
**Net loss for this user's goal, as things stand — stay on dense Qwen3.8-27B.** Reasoning:
|
||
|
||
- The user's target (500k×2 or 500k×4, on a $20-30/GB-VRAM budget, GPUs in 32GB increments) is a
|
||
**VRAM-budget-constrained** goal, and §11.7 shows Flash-Next's *weights alone* (111-354GB depending
|
||
on quant) dwarf the entire dense-model total-VRAM figure from §4/§11.7 (33-183GB depending on quant)
|
||
at every parallelism target. Flash-Next's much cheaper per-token KV cache (§11.5, real and verified)
|
||
is a rounding error next to that weight-size gap — the "2.67× cheaper KV" win doesn't come close to
|
||
offsetting a "6-20× larger weight footprint," so at $20-30/GB-VRAM the *dense* model reaches 500k×2
|
||
or 500k×4 for a fraction of the card count and dollar cost that Flash-Next needs even at its lowest
|
||
usable quant (`UD-Q4_K_XL`, 4-5 cards minimum) — before even factoring in §11.1's immaturity flags.
|
||
- The one scenario that could flip this verdict is `--n-cpu-moe` actually working well for this
|
||
architecture (§11.8) — if most of those 111-354GB of expert weights can sit in system RAM at
|
||
acceptable throughput, Flash-Next's GPU-VRAM number could shrink dramatically and its real KV-cache
|
||
advantage would start to matter. That is untested here and shouldn't be assumed; it's the one
|
||
concrete next step worth trying on the actual server before ruling Flash-Next out permanently.
|
||
- Independent of VRAM: §11.1's primary-source-flagged correctness caveats (unverified conv branch,
|
||
3%-divergence QSA, prefill-position-0-only PLE exactness) are a real quality/stability risk on a
|
||
production coding-agent stack that dense Qwen3.8-27B simply doesn't carry, since it's been running
|
||
in this repo already.
|
||
|
||
---
|
||
|
||
## 12. 4-parallel × 500k scenario — all four combinations side by side
|
||
|
||
Per §2's already-established, cited rule (`n_ctx_per_seq = n_ctx / n_parallel`,
|
||
[ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681)), the same division
|
||
applies at 4 slots: **500k tokens on each of 4 parallel slots requires `--ctx-size 2,000,000`** —
|
||
double the 2-parallel target's `--ctx-size 1,000,000`, for the same reason 500k×2 needed double
|
||
262,144. This isn't a new mechanism, just the same formula at `--parallel 4`.
|
||
|
||
Full per-quant-tier tables for all four combinations are in §11.7 above (dense×2, dense×4, Flash-Next×2,
|
||
Flash-Next×4 are each their own table there). Headline comparison at the KV precision already proven
|
||
in production in this repo (`q8_0`) and each model's respective current/cheapest-usable weight quant:
|
||
|
||
| Scenario | `--ctx-size` | Weight quant | q8_0-KV total VRAM | Cards (32GB) |
|
||
|---|---|---|---|---|
|
||
| Dense × 2 (or 1M stretch) | 1,000,000 | Q4_K_XL (17.6GB, current) | ~48-54GB | **2** |
|
||
| Dense × 4 | 2,000,000 | Q4_K_XL (17.6GB, current) | ~82-85GB | **3** |
|
||
| Flash-Next × 2 (or 1M stretch) | 1,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~126-129GB | **4**, edge |
|
||
| Flash-Next × 4 | 2,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~137-140GB | **5** |
|
||
|
||
**4-parallel × 500k reachability against this repo's existing roadmap stages (§9):**
|
||
|
||
- **(a) Current 1×R9700 32GB:** none of the four combinations fit — not even dense×2 at any weight/KV
|
||
quant (§4's own conclusion, unchanged).
|
||
- **(b) The 2-3×R9700 roadmap already proposed in §9 (64-96GB):** covers **dense×2 fully** (stage 3, as
|
||
already established) and **dense×4 at `q4_0` KV with Q4_K_XL or Q6_K_XL weights** (~51-62GB, fits in
|
||
64-96GB) — but **not** dense×4 at higher weight quants (Q8_0/BF16 need 3-6 cards depending on KV
|
||
precision, per §11.7's dense×4 table) and **not any Flash-Next scenario** (minimum is 4 cards/128GB
|
||
even at the cheapest usable quant and tightest KV).
|
||
- **(c) The full 4-6×R9700 stretch scenario** (`server-planing.md`'s original plan, 128-192GB pooled):
|
||
covers **dense×4 at every weight quant up to BF16** (worst case ~91GB at BF16+q4_0, well inside
|
||
128GB) and **Flash-Next×2 at `UD-Q4_K_XL`** (126-140GB, fits a 5-card/160GB build, tight on a 4-card/
|
||
128GB one) — but **not** Flash-Next×4 at any weight quant above `UD-Q4_K_XL`, and not Flash-Next at
|
||
`BF16` under any parallelism (needs 12-13 cards, an entirely different scale of build than anything
|
||
in this doc's roadmap).
|
||
|
||
---
|
||
|
||
## Sources
|
||
|
||
- [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) — `--ctx-size`, `--parallel`, `--cache-type-k/v`, `--kv-unified`, `--cache-reuse`, `--n-cpu-moe`, `--override-tensor` flag definitions
|
||
- [ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681) — real server log proving `n_ctx_per_seq = n_ctx / n_parallel`
|
||
- [ggml-org/llama.cpp#27742](https://github.com/ggml-org/llama.cpp/pull/27742) — "model: add Qwen3.8-Flash-Next (qwen4exp)", merged 2026-08-27; architecture details, `n_stream == 1` / `-kvu` multi-slot requirement, and the conv-branch/QSA-divergence/PLE-prefill correctness caveats
|
||
- [Qwen/Qwen3.8-27B config.json](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json)
|
||
- [Qwen/Qwen3.8-Flash-Next config.json](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/raw/main/config.json)
|
||
- [Qwen/Qwen3-4B-Instruct-2507 config.json](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json)
|
||
- [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main) — 17.6GB (Q4_K_XL), 25.3GB (Q6_K_XL), 29GB (Q8_0), 54.67GB (BF16) weight sizes
|
||
- [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main) — 111.4GB (UD-Q4_K_XL), 169GB (UD-Q6_K_XL), 188GB (Q8_0), 354GB (BF16) weight sizes, summed from each quant's per-file listing
|
||
- [ggml-org/llama.cpp docs/docker.md](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md) — `server-cuda`/`server-rocm` separate image tags
|
||
- [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html)
|
||
- [Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2)
|
||
- [PCWorld — Threadripper PRO launch, 128 PCIe lanes](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html)
|
||
- [NVIDIA — GeForce RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/)
|
||
- [Lowyat.net — RTX 3060 official 170W TDP](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/)
|
||
- [BuildMyServer — GTX 1080 180W/PCIe3.0/1×8-pin spec listing](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu)
|
||
- [MSI — GT 710 1GD5 LP spec page](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification)
|
||
- [overclock3d — AMD Radeon AI PRO R9700 $1,299 MSRP](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/)
|
||
- [pricehistory.app — R9700 street price tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm)
|
||
- [bestvaluegpu.com — RTX 3090 used price tracker, Sep 2026](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/)
|
||
- [gpupoet.com — RTX 3060 12GB used listings](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060)
|
||
- [Qwen Code docs — overview](https://qwenlm.github.io/qwen-code-docs/en/users/overview/)
|
||
- [Qwen Code docs — Auto Mode (fast-model pattern)](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/)
|
||
- This repo: [`docker-compose.yml`](../../docker-compose.yml), [`.env.example`](../../.env.example), [`docs/server-planing.md`](../server-planing.md), [`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), [`qwen3.8-27b-tool-calling.md`](qwen3.8-27b-tool-calling.md), [`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md), [`fast-model-choice.md`](fast-model-choice.md)
|
||
|
||
## Confidence/uncertainty summary
|
||
|
||
- **High confidence:** the KV-cache-per-token formulas for both dense models and Flash-Next (computed
|
||
directly from each model's own `config.json`, same method this repo's prior research already used
|
||
and cross-checked); the `n_ctx_per_seq = n_ctx / n_parallel` division behavior (directly evidenced by
|
||
a real server log in a llama.cpp GitHub issue, and independently already documented in this repo's
|
||
own `.env.example`) — and confirmed to apply identically at `--parallel 4` since the mechanism is
|
||
parallel-count-agnostic; official TDP figures for GTX 1080, RTX 2080, RTX 3060, RTX 3090, GT 710
|
||
(each cross-checked against 2+ independent spec listings or the vendor's own product page); the
|
||
TR4-is-PCIe3/sTRX4-is-PCIe4 generational split (direct launch-coverage quote); the existence of
|
||
separate `server-cuda`/`server-rocm` llama.cpp image tags; Qwen3.8-Flash-Next's `config.json`
|
||
architecture params and PR #27742's merge date/status and its own stated correctness caveats and
|
||
`-kvu` multi-slot requirement (all directly quoted from the primary source); the weight file sizes
|
||
for both models at all four quant tiers (summed directly from each HF repo's real file listing, not
|
||
estimated).
|
||
- **Medium confidence:** the compute-buffer/runtime-overhead estimate in §4/§11.7 (+3-6 GiB) —
|
||
extrapolated from one in-repo incident's before/after numbers, not a llama.cpp-documented formula,
|
||
and carried over to Flash-Next without re-derivation for its different architecture; the Gated
|
||
DeltaNet fixed recurrent-state size in §11.4 (~108 MiB/slot) — this doc's own derivation from the
|
||
published head-dimension config, not a value found in llama.cpp source or docs; whether
|
||
`--cache-type-k/v` quantization applies identically to Flash-Next's 12 full-attention layers as it
|
||
does to a plain transformer (reasonable extrapolation, not directly confirmed for this architecture);
|
||
whether `--n-cpu-moe`/`--override-tensor` actually work against Flash-Next's specific MoE tensor
|
||
layout (architecture-agnostic mechanism, but untested against this model by any primary source found);
|
||
real-world PCIe lane availability for 6 GPUs on a specific sTRX4 board (§5) — no single board's exact
|
||
lane map was fetched, this is a reasonable-but-unverified estimate from typical sTRX4 board behavior.
|
||
- **Low confidence / explicitly estimated, not cited fact:** all used-GPU marketplace pricing (§8) —
|
||
live asking-price snapshots from a single search pass, not sold-price data; the used sTRX4
|
||
CPU+motherboard combo price in the roadmap's stage 4 (§9) — not researched at all in this pass,
|
||
flagged as a placeholder estimate; whether YaRN-scaled 500k/1M context actually holds output
|
||
quality for either Qwen3.8-27B or Qwen3.8-Flash-Next — no primary source (Qwen's own docs included)
|
||
publishes long-context quality benchmarks past the 262,144 native length for either model, so this is
|
||
a known-unknown carried forward from each model card's "YaRN-extensible" claim, not a verified
|
||
capability; whether the specific `ghcr.io/ggml-org/llama.cpp:server-rocm` image currently cached on
|
||
this repo's server actually postdates PR #27742's 2026-08-27 merge — not checked against the live
|
||
server in this pass, flagged as an action item in §11.1 rather than a confirmed fact.
|