Files
LLM-Server/docs/research/500k-context-2-parallel-agents-gpu-roadmap.md
T

738 lines
56 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Research: hardware roadmap to 500k-token context × 2 parallel agents (1M stretch)
**Date:** 2026-09-11
**Question:** What VRAM does 500k-token context × 2 parallel llama-server slots (and a 1M-token
stretch goal) actually cost for the Qwen3 family, and what hardware roadmap gets there from the
current single-R9700 setup — given the user's stated plan to add an older (PCIe 4.0) Threadripper
for lane count, reuse existing RAM/PSU (~200W headroom / one spare 8-pin), mix in already-owned
NVIDIA cards (GTX 1080 8GB, RTX 2080 8GB, GT 710 1GB) for the classifier role, and price used GPUs
at roughly $20-30/GB VRAM?
**Answer, short version:** The two goals ("500k × 2 parallel" and "1M stretch") turn out to need
**the same total VRAM budget** — because of how llama-server's `--ctx-size` and `--parallel` interact
(§2), 500k × 2 slots and a single 1M-token slot both require setting `--ctx-size 1000000`. At
`q4_0`-quantized KV cache that's **~33 GB** (weights + KV) for Qwen3.8-27B, at `q8_0` it's **~48 GB**,
at fp16 it's **~79 GB** — before compute-buffer overhead. That does not fit on the current single
32GB R9700 at any KV precision, and comfortably fits on two 32GB-class cards only at `q8_0`/`q4_0`.
The user's $20-30/GB pricing intuition holds for last-gen used consumer cards (RTX 3060 12GB) but
**not** for RTX 3090 24GB (~$44/GB currently) or a second R9700 (~$41/GB, new — no used market yet
for a card released mid-2026). The stated Threadripper plan needs to specifically target the
**non-PRO Threadripper 3000 series on sTRX4** (64 lanes, PCIe 4.0) — older Threadripper on the
original TR4 socket (1000/2000 series) is PCIe 3.0 only, which doesn't match the user's own PCIe 4.0
requirement. The power budget (~200W / one spare 8-pin) is exhausted by a *single* mid-tier used GPU
addition — a PSU upgrade is not optional past the very first stage. See §7 for the roadmap.
---
## 1. Current state (from this repo)
From `docker-compose.yml` and `.env.example` at the repo root:
- **Main model:** `Qwen3.8-27B-UD-Q4_K_XL.gguf` (17.6 GB weights), `--ctx-size 262144`,
`--parallel 2`, `--flash-attn on`, `--cache-type-k q8_0 --cache-type-v q8_0`, `--n-gpu-layers 999`,
on one AMD Radeon AI PRO R9700 (32GB, ROCm/HIP, `gfx1201`).
- **Classifier ("fast") model:** `Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`, `--ctx-size 65536`,
`--parallel 1`, `--n-gpu-layers 28` (partial offload), `--cache-type-k/v q4_0`, its own container on
the *same* R9700, sharing VRAM with the main model — see
[`docker-compose.yml`](../../docker-compose.yml) lines ~58-99 and
[`fast-model-choice.md`](fast-model-choice.md).
- `.env.example` already documents the exact fact this research turns on: *"Each slot gets
`LLAMA_CTX_SIZE / LLAMA_PARALLEL` tokens of context"* — i.e. today's 262144 ctx-size ÷ 2 parallel
slots means each real request only gets **~131K tokens**, not the full 262144, confirmed in-repo
before any external source was checked.
- Prior research already worked out the KV-cache formula for this exact model
([`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md)) — this doc reuses and extends that math for the
500k/1M targets rather than re-deriving it.
`docs/server-planing.md` describes a **different, earlier plan**: a 4× AMD Radeon AI PRO R9700 rig
on a Gigabyte MZ32-AR0 (single-socket SP3/EPYC, 128 PCIe 4.0 lanes), fully AMD/ROCm. The user's plan
in this ticket is not that — it pivots toward an older **Threadripper** (SP3's sibling desktop-HEDT
socket family, not SP3 itself) and explicitly wants to mix in already-owned **NVIDIA** cards. These
two plans are **not the same build** and, per §6, ROCm and CUDA cards cannot share one llama.cpp
process — they can only coexist as separate containers on separate cards. Treat `server-planing.md`
as superseded context, not the active plan, unless the user says otherwise.
---
## 2. llama-server parallelism: does each slot get its own full `--ctx-size`, or is it divided?
**Divided.** This is the single fact that changes the whole budget by 2×, confirmed from three
independent primary sources:
1. **This repo's own `.env.example`** (quoted above) already documents it for the current deployment.
2. **llama.cpp's own server README**, fetched directly
([`tools/server/README.md`](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)):
`--ctx-size (-c)`: *"size of the prompt context (default: 0, 0 = loaded from model)"*;
`--parallel (-np)`: *"number of server slots (default: -1, -1 = auto)"* — the docs list these as
independent flags, but don't spell out the division themselves.
3. **A real user's server log**, quoted verbatim in
[ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681), is the actual proof:
running with `--ctx-size 327680 --parallel 6` produces `n_ctx = 327680`,
`n_ctx_per_seq = 54613` — i.e. `327680 / 6 ≈ 54613`. The reporter explicitly asked for a
`--ctx-size-per-seq`-style flag to *avoid* this division; no such flag exists as of the fetch date.
Practical consequence: **to get 500,000 usable tokens on each of 2 parallel slots, `--ctx-size` must
be set to 1,000,000, not 500,000.** The KV cache is sized off the *total* `--ctx-size`
(`--kv-unified`, on by default when slots are auto per the README's `-kvu` entry, uses one shared
pool sized to the full `n_ctx`) — so the VRAM cost of "500k × 2 parallel" and "one 1M-token slot"
is **identical**: both require `--ctx-size 1000000`. This is a genuinely useful finding for the
roadmap — reaching the 500k×2 target and the 1M stretch goal cost the same VRAM; the only difference
is `--parallel 1` vs `--parallel 2` at deploy time, a config change with zero extra hardware cost.
`--cache-type-k` / `--cache-type-v` accept `f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1`
(default `f16`), per the same README fetch. The repo already uses `q8_0` on the main model and `q4_0`
on the classifier, so both quantization tiers used in the math below are already-proven-working
configurations in this stack, not hypothetical flags.
---
## 3. KV-cache math per model
### Qwen3.8-27B (hybrid Gated-DeltaNet / attention)
Reusing the architecture params already pulled from
[`Qwen/Qwen3.8-27B/config.json`](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json) in
[`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), re-verified directly for this doc: `num_hidden_layers=64`,
`full_attention_interval=4`**16 of 64 layers are standard KV-caching attention**, the other 48 are
Gated DeltaNet linear-attention layers with a small, context-length-*independent* recurrent state
(tens of MB total, negligible next to the attention KV cache — ignored below).
`num_key_value_heads=4` (GQA), `head_dim=256`. Native context `max_position_embeddings=262144`
(YaRN-extensible to 1M per the model card — **both the 500k and 1M targets exceed native context and
require RoPE/YaRN scaling**, which is a real quality caveat, not just a memory one — Qwen has not
published independent long-context quality benchmarks past native length that this research found).
Per-token KV cache, fp16, both K and V, across the 16 full-attention layers:
```
16 layers × 2 (K+V) × 4 kv_heads × 256 head_dim × 2 bytes = 64 KiB/token
```
| Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 |
|---|---|---|---|
| 262,144 (current) | ~16.0 GiB | ~8.0 GiB | ~4.0 GiB |
| 500,000 | ~30.5 GiB | ~15.3 GiB | ~7.6 GiB |
| **1,000,000 (500k×2, or 1M stretch)** | **~61.0 GiB** | **~30.5 GiB** | **~15.3 GiB** |
(`q8_0` is 8-bit vs. fp16's 16-bit → exactly half; `q4_0` is 4-bit → exactly quarter, per llama.cpp's
own cache-type byte widths.)
### Qwen3-4B-Instruct-2507 (plain GQA transformer, classifier role)
From [`Qwen/Qwen3-4B-Instruct-2507/config.json`](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json)
(already pulled in [`fast-model-choice.md`](fast-model-choice.md)): `num_hidden_layers=36` — every
layer is standard attention here (no hybrid split), `num_key_value_heads=8`, `head_dim=128`.
```
36 layers × 2 (K+V) × 8 kv_heads × 128 head_dim × 2 bytes = 144 KiB/token
```
The classifier's real transcript ceiling is ~40-50K tokens (qwen-code's own
`MAX_TRANSCRIPT_MESSAGES=40` × `MAX_HISTORICAL_ACTION_CHARS=4000`, per `fast-model-choice.md` §"what
actually shipped") — nowhere near 500k/1M, so the classifier does **not** need to grow for this
roadmap; it stays exactly as deployed today, on its own small allocation. Per-token cost is included
here only because it feeds the "does the classifier's dedicated GPU need to change" question in §6.
---
## 4. Total VRAM budget: 500k × 2 parallel, and the 1M stretch
Weights: `Qwen3.8-27B-UD-Q4_K_XL.gguf` is **17.6 GB**, confirmed directly from the
[unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main)
(already verified in `qwen3.8-27b-quant.md`).
Per §2, both "500k × 2 parallel" and "1M stretch" require `--ctx-size 1000000` — same KV budget:
| KV precision | KV cache | + weights (17.6 GB) | + est. compute-buffer/runtime overhead* | **Realistic total** |
|---|---|---|---|---|
| fp16 (default) | 61.0 GiB | 78.6 GiB | +3-6 GiB | **~82-85 GB** |
| q8_0 (proven in this stack today) | 30.5 GiB | 48.1 GiB | +3-6 GiB | **~51-54 GB** |
| q4_0 (proven in this stack today, on the classifier) | 15.3 GiB | 32.9 GiB | +3-6 GiB | **~36-39 GB** |
\* *Estimate, not a cited figure* — llama.cpp's flash-attention compute buffer scales closer to
linear than the unfused-attention path, per this repo's own measured note in `docker-compose.yml`'s
`qwen-classifier` comment (unfused attention buffers ballooned unexpectedly at 65536 ctx; flash-attn
fixed it). `--flash-attn on` is already the deployed default for the main model, so the linear-ish
regime applies, but no primary source gives an exact formula for this buffer size at 1M context — the
+3-6 GiB band is this doc's estimate based on the ratio observed in that in-repo incident, not a
llama.cpp-documented number. Budget for the high end of that range when sizing hardware.
**Bottom line:** at `q4_0` KV (the most aggressive, already-proven-in-this-repo tier), 500k×2 /
1M needs **~36-39 GB** total VRAM for the 27B model alone. That does not fit one 32GB card at any
precision — it needs at least two 32GB-class cards, or one ≥40GB card. At `q8_0` (the precision this
repo already runs in production for quality reasons), budget **~51-54 GB** — two 32GB cards (64GB
pooled) clears this with room to spare; a single 48GB-class card would not.
(§11 below extends this table to higher weight-quant tiers — Q6_K_XL, Q8_0, BF16 — for users who want
better output quality than `Q4_K_XL`, and to a Flash-Next alternative architecture; see §11.6-§11.7.)
---
## 5. CPU/motherboard: which Threadripper generations give PCIe 4.0, and how many lanes for GPUs
AMD's own product/chipset pages, cross-checked against the launch reviews that quote them directly:
| Platform | Socket | PCIe generation | Total CPU-provided lanes |
|---|---|---|---|
| Threadripper 1000/2000 series ("1920X", "2950X", etc.) | **TR4** | **PCIe 3.0 only** | 60-64 |
| Threadripper 3000 series (3960X/3970X/3990X) | **sTRX4** | **PCIe 4.0** | 64 |
| Threadripper 7000 series (non-PRO) | sTR5 | PCIe 5.0 (48 lanes) + PCIe 4.0 (24-32 lanes) | ~72-80 |
| Threadripper PRO 3000WX/5000WX | sWRX8 | PCIe 4.0 | **128** |
| Threadripper PRO 7000WX | sTR5 (WRX90) | PCIe 5.0 (128 lanes) + a few PCIe 3.0 | **128** |
Sources: [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html)
(sWRX8/socket listing), corroborated by
[Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch coverage](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2)
(*"the 3rd Gen TR CPUs carry the same 64 PCIe lanes but double bandwidth by moving from Gen 3.0 to
Gen 4.0"* — explicit confirmation TR4/1000-2000-series is PCIe 3.0 while sTRX4/3000-series is PCIe
4.0), [PCWorld — Threadripper PRO launch](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html)
(*"128 PCIe lanes"* for PRO).
**This directly matters for the user's plan.** "An older Threadripper... for more PCIe lanes" is
ambiguous between two real, very different chips:
- **TR4 (1000/2000 series)** — cheapest used option, but **PCIe 3.0** — does not meet the user's own
stated PCIe 4.0 requirement, and PCIe 3.0 x8 per GPU roughly halves inter-GPU/host transfer
bandwidth (matters more for training/tensor-parallel than for llama.cpp's inference-time layer
splitting, but still a real downgrade vs. the R9700's native PCIe 5.0).
- **sTRX4 (3000 series, non-PRO)** — the correct "older Threadripper with PCIe 4.0" target: 64 lanes,
4-5 years old, real used-market availability, no PRO price premium.
- **Threadripper PRO (3000WX/5000WX)** doubles the lane count to 128 but at meaningfully higher used
cost (workstation-tier, lower volume, sWRX8 boards are pricier than sTRX4/TRX40 boards) — worth it
only if 6 full-bandwidth (x16) GPU slots are actually needed; at x8-per-card (adequate for inference)
64 lanes already covers 6 GPUs with lanes to spare for NVMe/chipset.
**Lane budget for 6 GPUs on sTRX4 (64 lanes), estimated (no vendor spec gives a topology this
specific — treat this bullet as an estimate):** typical sTRX4 boards reserve ~4 lanes for the
chipset uplink and commonly wire 1-2 M.2 slots directly to the CPU (4 lanes each) — so realistic
GPU-available lanes land around 44-52 of the 64, i.e. **6 GPUs at x8 electrical each (48 lanes) is
plausible but board-model-dependent**; x16-each for 6 cards is not possible on 64 lanes regardless of
board. x8 electrical is not a meaningful inference-speed penalty for llama.cpp (weights are loaded
once; the ongoing per-token traffic across PCIe is small compared to compute), so this is an
acceptable tradeoff, not a real bottleneck for this workload.
---
## 6. Power budget vs. the ~200W / one spare 8-pin headroom
Official/vendor TDPs:
| Card | TDP | Source |
|---|---|---|
| GTX 1080 (owned) | 180W, one 8-pin | [confirmed 180W, PCIe 3.0 x16, 1× 8-pin](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu) — spec matches NVIDIA's own launch figures reported across multiple outlets incl. Tom's Hardware |
| RTX 2080 (owned) | 215W | Cross-checked across gpuzoo/cputronic/notebookcheck spec pages, consistent at 215W |
| GT 710 (owned) | ~19W, **no external power connector** (slot power only) | [MSI/EVGA/Zotac GT 710 spec pages](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification) |
| RTX 3060 12GB (candidate purchase) | 170W, one 8-pin | [NVIDIA-confirmed 170W TDP, one 8-pin connector](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/) |
| RTX 3090 24GB (candidate purchase) | 350W, two 8-pin, [NVIDIA's own RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/) lists 350W and a 750W PSU minimum | NVIDIA official |
| R9700 32GB (already deployed / "more of the same") | 300W (per this repo's `server-planing.md`, consistent with AMD's own R9700 product page framing it as a 300W-class card) | in-repo prior research |
**Against the stated ~200W / one spare 8-pin budget:**
- Adding **one RTX 3060 12GB** (170W, one 8-pin) is the *only* candidate in this list that fits the
stated headroom as-is — it uses the one spare connector and stays under 200W.
- Adding the already-owned **GTX 1080** (180W) as the classifier's dedicated card also just barely
fits (180W ≤ 200W, one 8-pin) — this is a genuinely free option since the card is already owned and
its power draw is within budget, unlike every purchase candidate below.
- Adding the already-owned **RTX 2080** (215W) **exceeds** the stated 200W headroom by 15W — technically
over budget on paper, though real-world draw is usually a bit under rated TDP; flag it as marginal,
not safely fitting.
- Adding a **second R9700** (300W) or an **RTX 3090** (350W, needs two 8-pin — the user has only one
spare) both blow well past the current power budget on both watts and connector count.
- **The GT 710 draws no meaningful power (~19W, no PCIe power connector at all)** — it is free from a
power-budget standpoint regardless of what else is added.
**PSU upgrade trigger:** the very first stage that adds *any* GPU beyond a GTX 1080-class card (180W,
one 8-pin) or an RTX 3060 12GB (170W, one 8-pin) exhausts the stated headroom. Any stage that reaches
for a second 32GB-class card (R9700 or equivalent) or any 300W+ card **requires a PSU upgrade before
that stage**, not after — see the roadmap table in §7 for exactly which stage that is.
---
## 7. Mixed-GPU feasibility: ROCm + CUDA, and is the GT 710 usable at all
**ROCm and CUDA are different llama.cpp builds, but that's exactly the pattern already in this
repo.** `ghcr.io/ggml-org/llama.cpp` publishes both `server-rocm` and `server-cuda` as separate,
independently-built image tags (confirmed present on the [ggml-org container registry](https://github.com/orgs/ggml-org/packages/container/llama.cpp)
and documented in [`docs/docker.md`](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md)
*"server-cuda: Same as `server` but compiled with CUDA support"*, *"server-rocm: Same as `server`
but compiled with ROCm support"*). You cannot mix backends inside one process/container, but you
**can** run one `server-rocm` container pinned to the R9700 and a separate `server-cuda` container
pinned to an NVIDIA card, simultaneously, on the same host — this is architecturally identical to
today's `llama-server` + `qwen-classifier` two-container split in `docker-compose.yml`, just with a
different image tag for the NVIDIA-backed service and NVIDIA's container runtime (`nvidia-container-toolkit`
+ `--gpus` / device reservation, the CUDA-world equivalent of this repo's `/dev/kfd`+`/dev/dri`+
numeric-GID ROCm pattern documented in
[`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md)). None of that doc's ROCm-specific
findings (the `GPU_MAX_HW_QUEUES=1` MES firmware workaround, the numeric-GID `group_add` fix) apply to
an NVIDIA/CUDA container — those are ROCm-stack-specific bugs, not general multi-GPU-container issues.
**Is this an implicit AMD→NVIDIA rebuild, or additive?** Worth surfacing explicitly since the two
source plans conflict on this: `server-planing.md` is an AMD-only, ROCm-only 4×R9700 plan. This
ticket's plan is **additive/mixed** — keep the R9700 running the main model under ROCm, and bolt on
NVIDIA cards under CUDA for secondary roles (classifier, or a second inference GPU for the big model
if going the "more of the same type" route means buying NVIDIA instead of more R9700s). Both are
internally consistent, but they are different end-states — flag this choice back to the user rather
than assuming one.
**Splitting the *main* 27B model itself across mixed AMD+NVIDIA silicon in one process is not
possible** — llama.cpp's multi-GPU tensor-split only works within a single backend build. To use
both an R9700 and an NVIDIA card for the *same* model's layers, all the compute-hosting cards need to
be the same backend (all-ROCm or all-CUDA) in that one process. This is why §5's roadmap treats "add
GPU capacity to the main model" and "add a GPU for the classifier" as separable purchases with
different backend constraints, not a single mixed pool.
**Is the GT 710 usable for anything in this pipeline? No.** Reasoning:
- 1GB VRAM cannot hold any meaningful fraction of either model's weights (17.6 GB / 2.4-4.3 GB) —
even a handful of transformer layers at Q4 quantization exceeds 1GB.
- It's Kepler-generation silicon (192 CUDA cores, no tensor cores) — llama.cpp's CUDA backend
technically supports pre-Turing cards, but at this VRAM size there's nothing to usefully offload.
- It draws power from the PCIe slot only, no external connector — genuinely free to keep installed.
- **Plausible actual use: dedicate it as the box's display-output card**, so every compute-capable
GPU (R9700, and whichever NVIDIA cards get added) can be fully headless/compute-only with none of
their VRAM or a display output tied up driving a monitor — a real, if minor, use for it. This is
this doc's own inference from the spec facts above, not a claim found in any primary source.
---
## 8. GPU market pricing vs. the $20-30/GB assumption
| Card | VRAM | Backend | Current used-market price (estimate — see caveat) | $/GB |
|---|---|---|---|---|
| RTX 3060 12GB | 12GB | CUDA | ~$240-300 used (eBay listings, [gpupoet.com tracker](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060): *"from $239"*, [eBay live listings](https://www.ebay.com/shop/rtx-3060-12gb) averaging ~$488 asking but with a $239 floor) | **~$20-25/GB** — matches the stated assumption |
| RTX 3090 24GB | 24GB | CUDA | ~$1,010-1,050 used ([bestvaluegpu.com Sep 2026 tracker](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/), [xda-developers coverage](https://www.xda-developers.com/used-rtx-3090-still-best-for-local-ai-in-value/)) | **~$42-44/GB** — well above the stated assumption |
| R9700 32GB ("more of the same type") | 32GB | ROCm | **New only — $1,299 MSRP**, street price $1,400-1,585 as of this research ([overclock3d](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/), [pricehistory.app tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm)) — too recent a release (2026) for a used market to exist yet | **~$41-50/GB, and not a used-market price at all** |
**Caveat on all three price figures:** these are live marketplace asking-price snapshots pulled via
web search on 2026-09-11, not sold-price data or a vendor spec sheet — treat as directional, not
exact. eBay asking prices in particular run above realized sale prices.
**Correction to the user's stated assumption:** $20-30/GB is a good estimate specifically for
**last-generation mainstream used cards** (RTX 3060 12GB fits it almost exactly) but **not** for
high-VRAM flagship cards like the RTX 3090 (~1.5-2× that rate) or for "more of the same type" R9700
units, which aren't used-market at all yet and sit even higher per GB than the 3090. If the plan is
"cheapest path to more VRAM," multiple RTX 3060 12GB cards (or similar mid-tier used cards) beat one
RTX 3090 on $/GB, at the cost of needing more PCIe slots and more total wattage/connectors to reach
the same aggregate VRAM — which is exactly the tradeoff the Threadripper lane-count plan in §5 is
for.
---
## 9. Step-by-step roadmap
All "resulting max context" figures assume `--parallel 2` and the KV precision stated; per §2, the
`--ctx-size` value shown is the *total* (pre-division) value to pass to llama-server.
| Stage | Hardware change | Est. cost | Backend | Usable VRAM (main-model pool) | Max context @ parallel=2 (`q4_0` KV) | PSU upgrade triggered? |
|---|---|---|---|---|---|---|
| **0 (current)** | 1× R9700 32GB, in production | $0 | ROCm | 32GB (shared with classifier) | ~131K/slot today at `q8_0` KV (262144 total ÷ 2) | No |
| **1 — classifier isolation** | Move classifier onto the already-owned **GTX 1080** (180W, own container, `server-cuda`), freeing the R9700 entirely for the main model. Matches the existing dual-model pattern qwen-code's own docs describe (§10) and this repo's `qwen-classifier` service already implements, just on separate silicon instead of a shared card. | $0 (already owned) | ROCm (main) + CUDA (classifier) | R9700's full 32GB now available to the main model alone | ~262K/slot @ `q8_0` (unchanged ctx-size, no more classifier contention) | **No** — 180W GTX 1080 fits the stated ~200W/one-8-pin headroom |
| **2 — second big-model GPU** | Add **one more 32GB-class card** for the main model. Cheapest correct-backend option: a second R9700 (~$1,300-1,585 new, ROCm, same backend as the first — required if tensor-splitting one model across two cards) | ~$1,300-1,585 | ROCm | 64GB pooled | `--ctx-size 500000 --parallel 1` fits at `q4_0` (~33GB) or `q8_0` (~48GB, tight but fits in 64GB) — **not yet 500k×2** | **Yes** — 300W card, no spare 8-pin left after stage 1 |
| **3 — reach 500k × 2 / 1M stretch** | No further hardware if stage 2's 64GB pool is used with `--cache-type-k/v q4_0`: `--ctx-size 1000000 --parallel 2` needs ~33-39GB (§4), fits inside 64GB with real headroom for the compute buffer. If `q8_0` KV is required instead (this repo's current quality bar for the main model), the ~51-54GB need is tight-to-marginal on 64GB — a **third** 32GB card (~96GB pool) removes the risk. | $0 (reuses stage 2) or +$1,300-1,585 for a 3rd card if `q8_0` KV is required | ROCm | 64GB (q4_0 case) or 96GB (q8_0 case) | **500k×2 parallel achieved**, and the 1M stretch goal is the *same config* with `--parallel 1` instead of 2 (§2) | Already upgraded at stage 2 |
| **4 — optional CPU/lane platform swap** | Only needed if the plan is to keep scaling past 2-3 big cards, or to add several small used cards (RTX 3060 12GB) for extra headroom/throughput rather than raw ctx-size. Swap to **non-PRO Threadripper 3000-series (sTRX4)** — 64 PCIe 4.0 lanes, ~x8-per-slot for up to 6 GPUs (§5). Threadripper PRO 3000WX/5000WX (128 lanes) only if x16-per-card matters or 6+ full-bandwidth slots are wanted. | Used sTRX4 CPU+board: roughly $400-800 combined on the used market (not independently priced in this pass — **estimate**, not cited) | n/a (platform only) | n/a | n/a | Independent of GPU wattage — driven by whatever GPU count/wattage stage 5+ adds |
| **5+ — scale-out via small used cards** | Add RTX 3060 12GB units (~$20-25/GB, the assumption that actually holds, §8) instead of more 32GB flagship cards, once lane count (stage 4) supports it — useful for extra parallel slots / throughput beyond the 500k×2 target rather than for raising ctx-size further (500k×2/1M is already met at stage 3). | ~$240-300/card | CUDA (separate container per §6) | +12GB pooled per card, but on a *different backend* from the ROCm main model — usable for extra classifier/small-model capacity or a separate CUDA-backend llama-server instance, not as additional tensor-split VRAM for the ROCm main model | Unchanged for the main model; adds parallel capacity elsewhere | Yes, cumulative — each additional 170W card needs PSU headroom stage 2 already consumed |
**Where the existing dual-model pattern sits in this roadmap:** it's stage 1, and it's free. The
qwen-code docs pattern (main model + a small, always-resident, non-thinking fast/classifier model —
see §10) is already implemented in this repo; the only roadmap-relevant change is *which GPU* the
classifier sits on, moving it off the R9700 entirely onto an already-owned NVIDIA card frees the
R9700's full 32GB for the 500k×2/1M push instead of splitting it with the classifier as happens
today.
### 9.1 Upgrade path, as diagrams
Diagram form of the same §9 table and §11.9's dense-vs-Flash-Next call — nothing new is claimed here,
this is a visual index back into the cited sections above.
**Stage-by-stage hardware path** (PSU-upgrade triggers and target reached called out inline):
```mermaid
flowchart TD
S0["Stage 0 — today<br/>1x R9700 32GB, ROCm<br/>classifier shares the card<br/>$0"]
S1["Stage 1 — classifier isolation<br/>+ GTX 1080 (owned, 180W, CUDA)<br/>R9700 freed for main model<br/>$0 · PSU OK (180W fits ~200W headroom)"]
S2["Stage 2 — 2nd big-model GPU<br/>+1x R9700 32GB (ROCm)<br/>64GB pooled<br/>~$1,300-1,585 · PSU UPGRADE REQUIRED (300W, no 8-pin left)"]
S3q4["Stage 3a — q4_0 KV<br/>--ctx-size 1,000,000 --parallel 2<br/>~33-39GB, fits in 64GB<br/>$0 (reuses stage 2)"]
S3q8["Stage 3b — q8_0 KV (current prod quality)<br/>~51-54GB, tight on 64GB<br/>+1x R9700 -> 96GB removes risk<br/>+~$1,300-1,585"]
TARGET(["500k x2 parallel reached<br/>= 1M stretch goal, same VRAM<br/>(--parallel 1 vs 2 is a config flag, §2)"])
S4["Stage 4 — platform swap (optional)<br/>sTRX4 Threadripper 3000, 64 PCIe4 lanes<br/>only needed past 2-3 big cards<br/>~$400-800 (estimate, §9)"]
S5["Stage 5+ — scale out<br/>+RTX 3060 12GB cards (CUDA, separate backend)<br/>extra parallel/throughput, not more ctx-size<br/>~$240-300/card · PSU upgrade each card"]
S0 --> S1 --> S2
S2 --> S3q4 --> TARGET
S2 --> S3q8 --> TARGET
TARGET -.->|"only if scaling past this"| S4 --> S5
style TARGET fill:#2e7d32,color:#fff,stroke:#1b5e20
style S2 fill:#8a5a00,color:#fff,stroke:#5c3d00
style S5 fill:#8a5a00,color:#fff,stroke:#5c3d00
```
**Model choice, and the one open question that could change it** (§11.9):
```mermaid
flowchart TD
Q{"Goal: 500k-1M ctx<br/>within a $20-30/GB VRAM budget?"}
D["Dense Qwen3.8-27B<br/>17.6-54.7GB weights (Q4_K_XL-BF16)<br/>reaches 500k x2 on 2-3 cards,<br/>500k x4 on 2-6 cards depending on quant<br/>(§11.7 tables)"]
F{"Try --n-cpu-moe:<br/>offload MoE experts to system RAM?<br/>(untested for this model, §11.8)"}
FBAD["Flash-Next, all-GPU weights<br/>111-354GB just for weights<br/>needs 4-13 cards before any KV cost<br/>NOT recommended at this budget (§11.9)"]
FGOOD["Flash-Next, experts in system RAM<br/>GPU VRAM could shrink a lot<br/>2.67x cheaper KV/token becomes relevant<br/>UNVERIFIED — prototype on real server first"]
CAVEAT["+ real caveat either way:<br/>PR #27742 flags unverified conv branch,<br/>3% QSA divergence, prefill-pos-0-only PLE<br/>(§11.1) — dense model carries no such flag"]
Q --> D
Q -->|"considering Flash-Next instead"| F
F -->|"works well"| FGOOD
F -->|"doesn't help / untested"| FBAD
FGOOD --> CAVEAT
FBAD --> CAVEAT
style D fill:#2e7d32,color:#fff,stroke:#1b5e20
style FBAD fill:#8a1c1c,color:#fff,stroke:#5c1212
style FGOOD fill:#8a5a00,color:#fff,stroke:#5c3d00
```
---
## 10. Qwen-code's own docs on the fast-model/classifier pattern
Fetched directly per the user's link:
[qwenlm.github.io/qwen-code-docs/en/users/overview/](https://qwenlm.github.io/qwen-code-docs/en/users/overview/)
**the overview page itself does not describe the dual-model/classifier pattern**; it only covers
single-model-provider setup (Alibaba ModelStudio / third-party / custom provider), one model at a
time. The actual fast-model/classifier documentation lives on the **Auto Mode** page instead, which
this repo's own `fast-model-choice.md` already fetched and cited in detail:
[qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/) —
summary (see `fast-model-choice.md` §1 for the full quote): a two-stage classifier gate, both stages
using "your configured fast model (`/model --fast`)", Stage 1 a ~300ms `{shouldBlock}`-only check,
Stage 2 a ~3-5s chain-of-thought review that only runs on a Stage-1 block. Nothing in either page
gives a recommended *context size* or *model size* for the fast model beyond what's implied by that
latency budget — this repo's own prior research (`fast-model-choice.md`) derived the actual context
requirement from qwen-code's source code instead (`packages/core/src/permissions/classifier-transcript.ts`),
since the docs pages don't state one. No new information changes that prior doc's conclusion; this
section exists to confirm the overview page was checked directly as instructed and doesn't contradict
or add to it.
---
## 11. Alternative: Qwen3.8-Flash-Next (MoE, hybrid attention)
The user also wants to weigh switching (or adding) **Qwen3.8-Flash-Next** — a 125B-total/6B-active MoE
with a hybrid recurrent-attention architecture — against staying on dense Qwen3.8-27B, and separately
wants this section to cover **going up in weight quant** (Q4_K_XL → Q6_K_XL → Q8_0 → BF16/fp16) for
*both* models, not just Q4. Feasibility first, since it gates everything else.
### 11.1 Feasibility verdict: supported, but immature — read before trusting any number below
Checked directly against the primary sources the task named:
- **llama.cpp mainline support exists.** [PR #27742](https://github.com/ggml-org/llama.cpp/pull/27742)
("model: add Qwen3.8-Flash-Next (qwen4exp)") was **merged into `master` on 2026-08-27** by ngxson.
It adds the full architecture: Gated DeltaNet layers (sigmoid-gated linear attention), QSA
("Qwen Sparse Attention", operating at micro-block granularity), hyper-connections, and the PLE
n-gram embedding table. `llama.cpp`'s own docs list CPU/CUDA/Metal/ROCm as supported backends for
it — this is not a CUDA-only feature.
- **This repo's pinned image is a floating tag, not a version pin.** `docker-compose.yml` runs
`ghcr.io/ggml-org/llama.cpp:server-rocm` with no date/digest suffix — a rolling "latest ROCm server
build" tag, not a release version. The merge is from 2026-08-27, and today is 2026-09-11 (~2 weeks
later), so a **fresh pull** of `server-rocm` should include it — but whatever image is already
cached/running on the R9700 box may predate the merge. **Action before touching this model on the
server: `docker compose pull llama-server` and check the startup log's build/commit banner is dated
on/after 2026-08-27**, not just "the tag says server-rocm."
- **Real, primary-source-flagged immaturity — this is the part that should temper enthusiasm.** The
PR's own description/review discussion states: *"The conv branch itself is still numerically
unverified because the fixture zeroes its weights"*; QSA sparse attention *"diverges on 3 percent of
positions"* above its budget threshold; the PLE depthwise convolution *"is exact only for a prefill
that starts at position 0"* (i.e. correctness is not guaranteed once `--cache-reuse`/prompt-caching
is in play — a flag this repo already turns on for the dense model per the latest commit). None of
that is disqualifying, but it is a primary-source admission that this is a fresh, not-fully-verified
implementation, not a mature, widely-battle-tested one like the dense Qwen3.8-27B path.
- **Multi-slot serving needs an explicit new flag.** The same PR states: *"`set_input_qsa` asserted
`n_stream == 1`, so llama-server could not serve this model with more than one slot unless `-kvu`
was passed."* Per the server README (§2), `--kv-unified`/`-kvu` defaults to enabled **only when slot
count is auto** (`-1`). This repo's compose file sets `--parallel ${LLAMA_PARALLEL:-2}` **explicitly**
(not auto) — so adopting Flash-Next with `--parallel` > 1 requires **adding `--kv-unified` (or
`-kvu`) to the launch flags**, a real deploy-time change, not something that "just works" by copying
today's flag set onto a new model file.
**Verdict: yes, runnable** on this repo's backend (ROCm, mainline, no dev branch needed) as long as the
image is pulled after 2026-08-27 and `-kvu` is added for multi-slot use — but treat it as
**usable-with-caution**, not a drop-in swap, given the PR author's own unresolved-correctness notes.
### 11.2 Architecture, verified against `config.json` directly
Fetched from `Qwen/Qwen3.8-Flash-Next`'s `config.json` (unsloth's GGUF repo repackages the same base
model): `num_hidden_layers=48`, `hidden_size=2560`, `num_attention_heads=24`, `num_key_value_heads=2`,
`head_dim=256`, `max_position_embeddings=262144` (same native/extensible-to-1M framing as the dense
model — same YaRN quality caveat from §3 applies here too, unverified past native length), `num_experts=512`,
`num_experts_per_tok=10`, and the linear-attention head config: `linear_num_key_heads=16`,
`linear_num_value_heads=48`, `linear_key_head_dim=128`, `linear_value_head_dim=128`.
Layer pattern (confirmed both from the model card's own description and `config.json`'s
`full_attention_interval=4`): every 4th layer is full/QSA attention, the other 3 are Gated DeltaNet —
**12 of 48 layers grow a real KV cache; the other 36 have a fixed-size recurrent state that does not
grow with context length.** (24% full-attention layers vs. the dense model's 16-of-64 = 25% — similar
ratio, but the *absolute* per-layer KV cost differs because `num_key_value_heads` is 2 here vs. 4 on
the dense model — see below.)
### 11.3 Per-token growing-KV-cache cost
```
12 full-attention layers × 2 (K+V) × 2 kv_heads × 256 head_dim × 2 bytes (fp16) = 24 KiB/token
```
| Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 |
|---|---|---|---|
| 262,144 (native) | ~6.0 GiB | ~3.0 GiB | ~1.5 GiB |
| 500,000 | ~11.4 GiB | ~5.7 GiB | ~2.9 GiB |
| **1,000,000 (500k×2, or 1M stretch)** | **~22.9 GiB** | **~11.4 GiB** | **~5.7 GiB** |
| **2,000,000 (500k×4)** | **~45.8 GiB** | **~22.9 GiB** | **~11.4 GiB** |
`--cache-type-k/v` are the same generic llama.cpp KV-cache-quantization flags used elsewhere in this
doc; nothing in the PR or the server README suggests they're handled differently for the 12
full-attention layers of a hybrid model — they quantize the same growing K/V buffers as on a plain
transformer. (No primary source explicitly confirms this for *this* architecture specifically — flagged
as a reasonable extrapolation, not a directly-cited fact, same caveat class as this doc's other
estimates.)
### 11.4 Fixed (non-growing) recurrent state — Gated DeltaNet layers
The 36 Gated DeltaNet layers each keep a fixed-size recurrent state (an outer-product-style
key×value matrix per head) that does **not** scale with context length — only with slot/sequence
count. Sized from `config.json`'s linear-attention head params:
```
36 layers × linear_num_value_heads(48) × linear_key_head_dim(128) × linear_value_head_dim(128) × 4 bytes (fp32 state)
≈ 36 × 48 × 128 × 128 × 4 bytes ≈ 108 MiB per slot
```
This is **this doc's own derivation from the published head-dimension params, not a value pulled
directly from llama.cpp source or docs** — the PR text confirms the state exists per-stream/per-slot
but doesn't publish an exact byte formula, so treat the ~108 MiB/slot figure as an estimate, medium
confidence. Even at 4 parallel slots that's under half a gigabyte — **negligible** next to both the
growing KV cache (GBs) and the weights (tens to hundreds of GB) computed below. The headline
implication holds regardless of the exact multiplier: Flash-Next's "big memory line item" is the MoE
weights, not the attention state of any kind.
### 11.5 Magnitude vs. the dense model — how much cheaper is KV, really
At the same total ctx-size, Flash-Next's growing KV cache is **24 KiB/token vs. the dense model's
64 KiB/token — 2.67× smaller**, i.e. Flash-Next's KV budget is **37.5%** of the dense model's at
identical context length. This is a real, significant win *for the KV-cache line item specifically*
but see §11.9: it's a much smaller slice of a much bigger total, because the weights move the other
way by a far larger factor.
### 11.6 Weight sizes — verified from each unsloth GGUF repo's actual file listing
Fetched directly from the HF file trees (not estimated from ratios), current as of this research pass:
| Quant tier | Qwen3.8-27B (dense) | Qwen3.8-Flash-Next (MoE) |
|---|---|---|
| Q4_K_XL (`UD-Q4_K_XL`) | **17.6 GB** (existing baseline) | **111.4 GB** (4 parts: 10.9MB + 49.9GB + 49.4GB + 12.1GB) |
| Q6_K_XL (`UD-Q6_K_XL`) | **25.3 GB** | **169 GB** (6 parts) |
| Q8_0 | **29 GB** | **188 GB** (6 parts) |
| BF16/fp16 | **54.67 GB** (50GB + 4.67GB, 2 parts) | **354 GB** (8 parts) |
Sources: [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main)
and its `BF16/` subfolder; [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main)
and its `UD-Q4_K_XL/`, `UD-Q6_K_XL/`, `Q8_0/`, `BF16/` subfolders (per-file sizes summed). The
preliminary Q8_0 figure floated before this research pass (~192GB) was slightly high — the real
listing sums to **188 GB**; everything else in the preliminary list was accurate to within rounding.
**The weight-quant axis and the KV-cache-quant axis are independent knobs.** Raising weight quality
(Q4_K_XL → BF16) does not require raising `--cache-type-k/v` — the two flags are unrelated, and this
repo already proves that pattern works (`q8_0` KV cache is deployed today against `Q4_K_XL` weights).
A user chasing **maximum output quality** can run e.g. **BF16 weights + `q4_0` KV cache** — full-precision
weights for quality, still-compressed KV for context budget — or any other combination in the tables
below; nothing about picking a higher weight quant forces a matching KV precision.
### 11.7 Total VRAM: does it fit, across quant tiers and both parallelism targets
All totals = weights + growing KV cache + an estimated **+3-6 GB** compute-buffer/runtime overhead
(same estimate band as §4, carried over — not re-derived for this architecture; flagged medium
confidence there too). "Cards" = ceil(total ÷ 32GB), i.e. how many R9700-class 32GB cards it takes.
#### Dense Qwen3.8-27B — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`)
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|---|---|---|---|
| Q4_K_XL (17.6GB) | ~79-85GB (**3**) | ~48-54GB (**2**) | ~33-39GB (**2**) |
| Q6_K_XL (25.3GB) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) | ~44-47GB (**2**) |
| Q8_0 (29GB) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) | ~47-50GB (**2**) |
| BF16 (54.67GB) | ~119-122GB (**4**) | ~88-91GB (**3**) | ~73-76GB (**3**) |
#### Dense Qwen3.8-27B — 500k × 4 parallel (`--ctx-size 2,000,000`)
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|---|---|---|---|
| Q4_K_XL (17.6GB) | ~143-146GB (**5**) | ~82-85GB (**3**) | ~51-54GB (**2**) |
| Q6_K_XL (25.3GB) | ~150-153GB (**5**) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) |
| Q8_0 (29GB) | ~154-157GB (**5**) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) |
| BF16 (54.67GB) | ~180-183GB (**6**) | ~119-122GB (**4**) | ~88-91GB (**3**) |
#### Qwen3.8-Flash-Next — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`)
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|---|---|---|---|
| UD-Q4_K_XL (111.4GB) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) | ~120-123GB (**4**) |
| UD-Q6_K_XL (169GB) | ~195-198GB (**7**) | ~183-186GB (**6**) | ~178-181GB (**6**) |
| Q8_0 (188GB) | ~214-217GB (**7**) | ~202-205GB (**7**) | ~197-200GB (**7**) |
| BF16 (354GB) | ~357-360GB (**12**) | ~357-360GB (**12**) | ~357-360GB (**12**) |
#### Qwen3.8-Flash-Next — 500k × 4 parallel (`--ctx-size 2,000,000`)
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|---|---|---|---|
| UD-Q4_K_XL (111.4GB) | ~160-163GB (**6**) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) |
| UD-Q6_K_XL (169GB) | ~218-221GB (**7**) | ~195-198GB (**7**) | ~183-186GB (**6**) |
| Q8_0 (188GB) | ~234-237GB (**8**) | ~211-214GB (**7**) | ~199-202GB (**7**) |
| BF16 (354GB) | ~397-400GB (**13**) | ~377-380GB (**12**) | ~366-369GB (**12**) |
(Flash-Next's KV precision barely moves the total at any weight quant above `UD-Q6_K_XL` — the weights
so dominate the budget that KV quantization stops mattering for the "how many cards" question. This
is the clearest signal in this whole section: for Flash-Next, the weight-quant choice is the entire
hardware-sizing decision; for the dense model, KV precision still matters a lot.)
### 11.8 CPU MoE-expert offload — the one lever that could change this calculus
Flash-Next is a 512-expert/10-active-per-token MoE, and llama.cpp has a purpose-built flag for exactly
this shape of model, confirmed directly from the server README: **`--n-cpu-moe`** — *"keep the Mixture
of Experts (MoE) weights of the first N layers in the CPU"* — plus the more general
**`--override-tensor`** (*"override tensor buffer type"*, pattern-matched by tensor name) that the same
flag is built on top of. Both are generic, architecture-agnostic llama.cpp mechanisms (they match on
tensor name patterns, not model type), so there's no reason to expect them not to apply to Flash-Next's
MoE tensors specifically — but this pass found **no primary source that has actually tested
`--n-cpu-moe` against this specific qwen4exp architecture**, so treat "it works here" as plausible,
not confirmed.
If it does work as expected, this changes the whole weight-VRAM picture in §11.7: the ~90-95% of
Flash-Next's weight footprint that's MoE expert tensors could live in system RAM while attention
projections, the shared/non-expert tensors, and the full KV cache stay on GPU — meaning a much smaller
GPU-VRAM number than the "all weights on GPU" tables above, at the cost of PCIe/RAM-bandwidth-bound
inference speed for whichever experts get selected per token (this repo has no benchmark of that
tradeoff, and it's highly system-RAM-bandwidth-dependent, so no number is given here — flagged as an
escape hatch worth prototyping directly on the server, not something this research values responsibly
without a real test run).
### 11.9 Net recommendation: dense Qwen3.8-27B vs. Flash-Next, for this user's stated goal
**Net loss for this user's goal, as things stand — stay on dense Qwen3.8-27B.** Reasoning:
- The user's target (500k×2 or 500k×4, on a $20-30/GB-VRAM budget, GPUs in 32GB increments) is a
**VRAM-budget-constrained** goal, and §11.7 shows Flash-Next's *weights alone* (111-354GB depending
on quant) dwarf the entire dense-model total-VRAM figure from §4/§11.7 (33-183GB depending on quant)
at every parallelism target. Flash-Next's much cheaper per-token KV cache (§11.5, real and verified)
is a rounding error next to that weight-size gap — the "2.67× cheaper KV" win doesn't come close to
offsetting a "6-20× larger weight footprint," so at $20-30/GB-VRAM the *dense* model reaches 500k×2
or 500k×4 for a fraction of the card count and dollar cost that Flash-Next needs even at its lowest
usable quant (`UD-Q4_K_XL`, 4-5 cards minimum) — before even factoring in §11.1's immaturity flags.
- The one scenario that could flip this verdict is `--n-cpu-moe` actually working well for this
architecture (§11.8) — if most of those 111-354GB of expert weights can sit in system RAM at
acceptable throughput, Flash-Next's GPU-VRAM number could shrink dramatically and its real KV-cache
advantage would start to matter. That is untested here and shouldn't be assumed; it's the one
concrete next step worth trying on the actual server before ruling Flash-Next out permanently.
- Independent of VRAM: §11.1's primary-source-flagged correctness caveats (unverified conv branch,
3%-divergence QSA, prefill-position-0-only PLE exactness) are a real quality/stability risk on a
production coding-agent stack that dense Qwen3.8-27B simply doesn't carry, since it's been running
in this repo already.
---
## 12. 4-parallel × 500k scenario — all four combinations side by side
Per §2's already-established, cited rule (`n_ctx_per_seq = n_ctx / n_parallel`,
[ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681)), the same division
applies at 4 slots: **500k tokens on each of 4 parallel slots requires `--ctx-size 2,000,000`**
double the 2-parallel target's `--ctx-size 1,000,000`, for the same reason 500k×2 needed double
262,144. This isn't a new mechanism, just the same formula at `--parallel 4`.
Full per-quant-tier tables for all four combinations are in §11.7 above (dense×2, dense×4, Flash-Next×2,
Flash-Next×4 are each their own table there). Headline comparison at the KV precision already proven
in production in this repo (`q8_0`) and each model's respective current/cheapest-usable weight quant:
| Scenario | `--ctx-size` | Weight quant | q8_0-KV total VRAM | Cards (32GB) |
|---|---|---|---|---|
| Dense × 2 (or 1M stretch) | 1,000,000 | Q4_K_XL (17.6GB, current) | ~48-54GB | **2** |
| Dense × 4 | 2,000,000 | Q4_K_XL (17.6GB, current) | ~82-85GB | **3** |
| Flash-Next × 2 (or 1M stretch) | 1,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~126-129GB | **4**, edge |
| Flash-Next × 4 | 2,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~137-140GB | **5** |
**4-parallel × 500k reachability against this repo's existing roadmap stages (§9):**
- **(a) Current 1×R9700 32GB:** none of the four combinations fit — not even dense×2 at any weight/KV
quant (§4's own conclusion, unchanged).
- **(b) The 2-3×R9700 roadmap already proposed in §9 (64-96GB):** covers **dense×2 fully** (stage 3, as
already established) and **dense×4 at `q4_0` KV with Q4_K_XL or Q6_K_XL weights** (~51-62GB, fits in
64-96GB) — but **not** dense×4 at higher weight quants (Q8_0/BF16 need 3-6 cards depending on KV
precision, per §11.7's dense×4 table) and **not any Flash-Next scenario** (minimum is 4 cards/128GB
even at the cheapest usable quant and tightest KV).
- **(c) The full 4-6×R9700 stretch scenario** (`server-planing.md`'s original plan, 128-192GB pooled):
covers **dense×4 at every weight quant up to BF16** (worst case ~91GB at BF16+q4_0, well inside
128GB) and **Flash-Next×2 at `UD-Q4_K_XL`** (126-140GB, fits a 5-card/160GB build, tight on a 4-card/
128GB one) — but **not** Flash-Next×4 at any weight quant above `UD-Q4_K_XL`, and not Flash-Next at
`BF16` under any parallelism (needs 12-13 cards, an entirely different scale of build than anything
in this doc's roadmap).
---
## Sources
- [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) — `--ctx-size`, `--parallel`, `--cache-type-k/v`, `--kv-unified`, `--cache-reuse`, `--n-cpu-moe`, `--override-tensor` flag definitions
- [ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681) — real server log proving `n_ctx_per_seq = n_ctx / n_parallel`
- [ggml-org/llama.cpp#27742](https://github.com/ggml-org/llama.cpp/pull/27742) — "model: add Qwen3.8-Flash-Next (qwen4exp)", merged 2026-08-27; architecture details, `n_stream == 1` / `-kvu` multi-slot requirement, and the conv-branch/QSA-divergence/PLE-prefill correctness caveats
- [Qwen/Qwen3.8-27B config.json](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json)
- [Qwen/Qwen3.8-Flash-Next config.json](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/raw/main/config.json)
- [Qwen/Qwen3-4B-Instruct-2507 config.json](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json)
- [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main) — 17.6GB (Q4_K_XL), 25.3GB (Q6_K_XL), 29GB (Q8_0), 54.67GB (BF16) weight sizes
- [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main) — 111.4GB (UD-Q4_K_XL), 169GB (UD-Q6_K_XL), 188GB (Q8_0), 354GB (BF16) weight sizes, summed from each quant's per-file listing
- [ggml-org/llama.cpp docs/docker.md](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md) — `server-cuda`/`server-rocm` separate image tags
- [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html)
- [Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2)
- [PCWorld — Threadripper PRO launch, 128 PCIe lanes](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html)
- [NVIDIA — GeForce RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/)
- [Lowyat.net — RTX 3060 official 170W TDP](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/)
- [BuildMyServer — GTX 1080 180W/PCIe3.0/1×8-pin spec listing](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu)
- [MSI — GT 710 1GD5 LP spec page](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification)
- [overclock3d — AMD Radeon AI PRO R9700 $1,299 MSRP](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/)
- [pricehistory.app — R9700 street price tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm)
- [bestvaluegpu.com — RTX 3090 used price tracker, Sep 2026](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/)
- [gpupoet.com — RTX 3060 12GB used listings](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060)
- [Qwen Code docs — overview](https://qwenlm.github.io/qwen-code-docs/en/users/overview/)
- [Qwen Code docs — Auto Mode (fast-model pattern)](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/)
- This repo: [`docker-compose.yml`](../../docker-compose.yml), [`.env.example`](../../.env.example), [`docs/server-planing.md`](../server-planing.md), [`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), [`qwen3.8-27b-tool-calling.md`](qwen3.8-27b-tool-calling.md), [`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md), [`fast-model-choice.md`](fast-model-choice.md)
## Confidence/uncertainty summary
- **High confidence:** the KV-cache-per-token formulas for both dense models and Flash-Next (computed
directly from each model's own `config.json`, same method this repo's prior research already used
and cross-checked); the `n_ctx_per_seq = n_ctx / n_parallel` division behavior (directly evidenced by
a real server log in a llama.cpp GitHub issue, and independently already documented in this repo's
own `.env.example`) — and confirmed to apply identically at `--parallel 4` since the mechanism is
parallel-count-agnostic; official TDP figures for GTX 1080, RTX 2080, RTX 3060, RTX 3090, GT 710
(each cross-checked against 2+ independent spec listings or the vendor's own product page); the
TR4-is-PCIe3/sTRX4-is-PCIe4 generational split (direct launch-coverage quote); the existence of
separate `server-cuda`/`server-rocm` llama.cpp image tags; Qwen3.8-Flash-Next's `config.json`
architecture params and PR #27742's merge date/status and its own stated correctness caveats and
`-kvu` multi-slot requirement (all directly quoted from the primary source); the weight file sizes
for both models at all four quant tiers (summed directly from each HF repo's real file listing, not
estimated).
- **Medium confidence:** the compute-buffer/runtime-overhead estimate in §4/§11.7 (+3-6 GiB) —
extrapolated from one in-repo incident's before/after numbers, not a llama.cpp-documented formula,
and carried over to Flash-Next without re-derivation for its different architecture; the Gated
DeltaNet fixed recurrent-state size in §11.4 (~108 MiB/slot) — this doc's own derivation from the
published head-dimension config, not a value found in llama.cpp source or docs; whether
`--cache-type-k/v` quantization applies identically to Flash-Next's 12 full-attention layers as it
does to a plain transformer (reasonable extrapolation, not directly confirmed for this architecture);
whether `--n-cpu-moe`/`--override-tensor` actually work against Flash-Next's specific MoE tensor
layout (architecture-agnostic mechanism, but untested against this model by any primary source found);
real-world PCIe lane availability for 6 GPUs on a specific sTRX4 board (§5) — no single board's exact
lane map was fetched, this is a reasonable-but-unverified estimate from typical sTRX4 board behavior.
- **Low confidence / explicitly estimated, not cited fact:** all used-GPU marketplace pricing (§8) —
live asking-price snapshots from a single search pass, not sold-price data; the used sTRX4
CPU+motherboard combo price in the roadmap's stage 4 (§9) — not researched at all in this pass,
flagged as a placeholder estimate; whether YaRN-scaled 500k/1M context actually holds output
quality for either Qwen3.8-27B or Qwen3.8-Flash-Next — no primary source (Qwen's own docs included)
publishes long-context quality benchmarks past the 262,144 native length for either model, so this is
a known-unknown carried forward from each model card's "YaRN-extensible" claim, not a verified
capability; whether the specific `ghcr.io/ggml-org/llama.cpp:server-rocm` image currently cached on
this repo's server actually postdates PR #27742's 2026-08-27 merge — not checked against the live
server in this pass, flagged as an action item in §11.1 rather than a confirmed fact.