# Research: hardware roadmap to 500k-token context × 2 parallel agents (1M stretch) **Date:** 2026-09-11 **Question:** What VRAM does 500k-token context × 2 parallel llama-server slots (and a 1M-token stretch goal) actually cost for the Qwen3 family, and what hardware roadmap gets there from the current single-R9700 setup — given the user's stated plan to add an older (PCIe 4.0) Threadripper for lane count, reuse existing RAM/PSU (~200W headroom / one spare 8-pin), mix in already-owned NVIDIA cards (GTX 1080 8GB, RTX 2080 8GB, GT 710 1GB) for the classifier role, and price used GPUs at roughly $20-30/GB VRAM? **Answer, short version:** The two goals ("500k × 2 parallel" and "1M stretch") turn out to need **the same total VRAM budget** — because of how llama-server's `--ctx-size` and `--parallel` interact (§2), 500k × 2 slots and a single 1M-token slot both require setting `--ctx-size 1000000`. At `q4_0`-quantized KV cache that's **~33 GB** (weights + KV) for Qwen3.8-27B, at `q8_0` it's **~48 GB**, at fp16 it's **~79 GB** — before compute-buffer overhead. That does not fit on the current single 32GB R9700 at any KV precision, and comfortably fits on two 32GB-class cards only at `q8_0`/`q4_0`. The user's $20-30/GB pricing intuition holds for last-gen used consumer cards (RTX 3060 12GB) but **not** for RTX 3090 24GB (~$44/GB currently) or a second R9700 (~$41/GB, new — no used market yet for a card released mid-2026). The stated Threadripper plan needs to specifically target the **non-PRO Threadripper 3000 series on sTRX4** (64 lanes, PCIe 4.0) — older Threadripper on the original TR4 socket (1000/2000 series) is PCIe 3.0 only, which doesn't match the user's own PCIe 4.0 requirement. The power budget (~200W / one spare 8-pin) is exhausted by a *single* mid-tier used GPU addition — a PSU upgrade is not optional past the very first stage. See §7 for the roadmap. --- ## 1. Current state (from this repo) From `docker-compose.yml` and `.env.example` at the repo root: - **Main model:** `Qwen3.8-27B-UD-Q4_K_XL.gguf` (17.6 GB weights), `--ctx-size 262144`, `--parallel 2`, `--flash-attn on`, `--cache-type-k q8_0 --cache-type-v q8_0`, `--n-gpu-layers 999`, on one AMD Radeon AI PRO R9700 (32GB, ROCm/HIP, `gfx1201`). - **Classifier ("fast") model:** `Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`, `--ctx-size 65536`, `--parallel 1`, `--n-gpu-layers 28` (partial offload), `--cache-type-k/v q4_0`, its own container on the *same* R9700, sharing VRAM with the main model — see [`docker-compose.yml`](../../docker-compose.yml) lines ~58-99 and [`fast-model-choice.md`](fast-model-choice.md). - `.env.example` already documents the exact fact this research turns on: *"Each slot gets `LLAMA_CTX_SIZE / LLAMA_PARALLEL` tokens of context"* — i.e. today's 262144 ctx-size ÷ 2 parallel slots means each real request only gets **~131K tokens**, not the full 262144, confirmed in-repo before any external source was checked. - Prior research already worked out the KV-cache formula for this exact model ([`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md)) — this doc reuses and extends that math for the 500k/1M targets rather than re-deriving it. `docs/server-planing.md` describes a **different, earlier plan**: a 4× AMD Radeon AI PRO R9700 rig on a Gigabyte MZ32-AR0 (single-socket SP3/EPYC, 128 PCIe 4.0 lanes), fully AMD/ROCm. The user's plan in this ticket is not that — it pivots toward an older **Threadripper** (SP3's sibling desktop-HEDT socket family, not SP3 itself) and explicitly wants to mix in already-owned **NVIDIA** cards. These two plans are **not the same build** and, per §6, ROCm and CUDA cards cannot share one llama.cpp process — they can only coexist as separate containers on separate cards. Treat `server-planing.md` as superseded context, not the active plan, unless the user says otherwise. --- ## 2. llama-server parallelism: does each slot get its own full `--ctx-size`, or is it divided? **Divided.** This is the single fact that changes the whole budget by 2×, confirmed from three independent primary sources: 1. **This repo's own `.env.example`** (quoted above) already documents it for the current deployment. 2. **llama.cpp's own server README**, fetched directly ([`tools/server/README.md`](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)): `--ctx-size (-c)`: *"size of the prompt context (default: 0, 0 = loaded from model)"*; `--parallel (-np)`: *"number of server slots (default: -1, -1 = auto)"* — the docs list these as independent flags, but don't spell out the division themselves. 3. **A real user's server log**, quoted verbatim in [ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681), is the actual proof: running with `--ctx-size 327680 --parallel 6` produces `n_ctx = 327680`, `n_ctx_per_seq = 54613` — i.e. `327680 / 6 ≈ 54613`. The reporter explicitly asked for a `--ctx-size-per-seq`-style flag to *avoid* this division; no such flag exists as of the fetch date. Practical consequence: **to get 500,000 usable tokens on each of 2 parallel slots, `--ctx-size` must be set to 1,000,000, not 500,000.** The KV cache is sized off the *total* `--ctx-size` (`--kv-unified`, on by default when slots are auto per the README's `-kvu` entry, uses one shared pool sized to the full `n_ctx`) — so the VRAM cost of "500k × 2 parallel" and "one 1M-token slot" is **identical**: both require `--ctx-size 1000000`. This is a genuinely useful finding for the roadmap — reaching the 500k×2 target and the 1M stretch goal cost the same VRAM; the only difference is `--parallel 1` vs `--parallel 2` at deploy time, a config change with zero extra hardware cost. `--cache-type-k` / `--cache-type-v` accept `f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1` (default `f16`), per the same README fetch. The repo already uses `q8_0` on the main model and `q4_0` on the classifier, so both quantization tiers used in the math below are already-proven-working configurations in this stack, not hypothetical flags. --- ## 3. KV-cache math per model ### Qwen3.8-27B (hybrid Gated-DeltaNet / attention) Reusing the architecture params already pulled from [`Qwen/Qwen3.8-27B/config.json`](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json) in [`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), re-verified directly for this doc: `num_hidden_layers=64`, `full_attention_interval=4` → **16 of 64 layers are standard KV-caching attention**, the other 48 are Gated DeltaNet linear-attention layers with a small, context-length-*independent* recurrent state (tens of MB total, negligible next to the attention KV cache — ignored below). `num_key_value_heads=4` (GQA), `head_dim=256`. Native context `max_position_embeddings=262144` (YaRN-extensible to 1M per the model card — **both the 500k and 1M targets exceed native context and require RoPE/YaRN scaling**, which is a real quality caveat, not just a memory one — Qwen has not published independent long-context quality benchmarks past native length that this research found). Per-token KV cache, fp16, both K and V, across the 16 full-attention layers: ``` 16 layers × 2 (K+V) × 4 kv_heads × 256 head_dim × 2 bytes = 64 KiB/token ``` | Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 | |---|---|---|---| | 262,144 (current) | ~16.0 GiB | ~8.0 GiB | ~4.0 GiB | | 500,000 | ~30.5 GiB | ~15.3 GiB | ~7.6 GiB | | **1,000,000 (500k×2, or 1M stretch)** | **~61.0 GiB** | **~30.5 GiB** | **~15.3 GiB** | (`q8_0` is 8-bit vs. fp16's 16-bit → exactly half; `q4_0` is 4-bit → exactly quarter, per llama.cpp's own cache-type byte widths.) ### Qwen3-4B-Instruct-2507 (plain GQA transformer, classifier role) From [`Qwen/Qwen3-4B-Instruct-2507/config.json`](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json) (already pulled in [`fast-model-choice.md`](fast-model-choice.md)): `num_hidden_layers=36` — every layer is standard attention here (no hybrid split), `num_key_value_heads=8`, `head_dim=128`. ``` 36 layers × 2 (K+V) × 8 kv_heads × 128 head_dim × 2 bytes = 144 KiB/token ``` The classifier's real transcript ceiling is ~40-50K tokens (qwen-code's own `MAX_TRANSCRIPT_MESSAGES=40` × `MAX_HISTORICAL_ACTION_CHARS=4000`, per `fast-model-choice.md` §"what actually shipped") — nowhere near 500k/1M, so the classifier does **not** need to grow for this roadmap; it stays exactly as deployed today, on its own small allocation. Per-token cost is included here only because it feeds the "does the classifier's dedicated GPU need to change" question in §6. --- ## 4. Total VRAM budget: 500k × 2 parallel, and the 1M stretch Weights: `Qwen3.8-27B-UD-Q4_K_XL.gguf` is **17.6 GB**, confirmed directly from the [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main) (already verified in `qwen3.8-27b-quant.md`). Per §2, both "500k × 2 parallel" and "1M stretch" require `--ctx-size 1000000` — same KV budget: | KV precision | KV cache | + weights (17.6 GB) | + est. compute-buffer/runtime overhead* | **Realistic total** | |---|---|---|---|---| | fp16 (default) | 61.0 GiB | 78.6 GiB | +3-6 GiB | **~82-85 GB** | | q8_0 (proven in this stack today) | 30.5 GiB | 48.1 GiB | +3-6 GiB | **~51-54 GB** | | q4_0 (proven in this stack today, on the classifier) | 15.3 GiB | 32.9 GiB | +3-6 GiB | **~36-39 GB** | \* *Estimate, not a cited figure* — llama.cpp's flash-attention compute buffer scales closer to linear than the unfused-attention path, per this repo's own measured note in `docker-compose.yml`'s `qwen-classifier` comment (unfused attention buffers ballooned unexpectedly at 65536 ctx; flash-attn fixed it). `--flash-attn on` is already the deployed default for the main model, so the linear-ish regime applies, but no primary source gives an exact formula for this buffer size at 1M context — the +3-6 GiB band is this doc's estimate based on the ratio observed in that in-repo incident, not a llama.cpp-documented number. Budget for the high end of that range when sizing hardware. **Bottom line:** at `q4_0` KV (the most aggressive, already-proven-in-this-repo tier), 500k×2 / 1M needs **~36-39 GB** total VRAM for the 27B model alone. That does not fit one 32GB card at any precision — it needs at least two 32GB-class cards, or one ≥40GB card. At `q8_0` (the precision this repo already runs in production for quality reasons), budget **~51-54 GB** — two 32GB cards (64GB pooled) clears this with room to spare; a single 48GB-class card would not. (§11 below extends this table to higher weight-quant tiers — Q6_K_XL, Q8_0, BF16 — for users who want better output quality than `Q4_K_XL`, and to a Flash-Next alternative architecture; see §11.6-§11.7.) --- ## 5. CPU/motherboard: which Threadripper generations give PCIe 4.0, and how many lanes for GPUs AMD's own product/chipset pages, cross-checked against the launch reviews that quote them directly: | Platform | Socket | PCIe generation | Total CPU-provided lanes | |---|---|---|---| | Threadripper 1000/2000 series ("1920X", "2950X", etc.) | **TR4** | **PCIe 3.0 only** | 60-64 | | Threadripper 3000 series (3960X/3970X/3990X) | **sTRX4** | **PCIe 4.0** | 64 | | Threadripper 7000 series (non-PRO) | sTR5 | PCIe 5.0 (48 lanes) + PCIe 4.0 (24-32 lanes) | ~72-80 | | Threadripper PRO 3000WX/5000WX | sWRX8 | PCIe 4.0 | **128** | | Threadripper PRO 7000WX | sTR5 (WRX90) | PCIe 5.0 (128 lanes) + a few PCIe 3.0 | **128** | Sources: [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html) (sWRX8/socket listing), corroborated by [Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch coverage](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2) (*"the 3rd Gen TR CPUs carry the same 64 PCIe lanes but double bandwidth by moving from Gen 3.0 to Gen 4.0"* — explicit confirmation TR4/1000-2000-series is PCIe 3.0 while sTRX4/3000-series is PCIe 4.0), [PCWorld — Threadripper PRO launch](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html) (*"128 PCIe lanes"* for PRO). **This directly matters for the user's plan.** "An older Threadripper... for more PCIe lanes" is ambiguous between two real, very different chips: - **TR4 (1000/2000 series)** — cheapest used option, but **PCIe 3.0** — does not meet the user's own stated PCIe 4.0 requirement, and PCIe 3.0 x8 per GPU roughly halves inter-GPU/host transfer bandwidth (matters more for training/tensor-parallel than for llama.cpp's inference-time layer splitting, but still a real downgrade vs. the R9700's native PCIe 5.0). - **sTRX4 (3000 series, non-PRO)** — the correct "older Threadripper with PCIe 4.0" target: 64 lanes, 4-5 years old, real used-market availability, no PRO price premium. - **Threadripper PRO (3000WX/5000WX)** doubles the lane count to 128 but at meaningfully higher used cost (workstation-tier, lower volume, sWRX8 boards are pricier than sTRX4/TRX40 boards) — worth it only if 6 full-bandwidth (x16) GPU slots are actually needed; at x8-per-card (adequate for inference) 64 lanes already covers 6 GPUs with lanes to spare for NVMe/chipset. **Lane budget for 6 GPUs on sTRX4 (64 lanes), estimated (no vendor spec gives a topology this specific — treat this bullet as an estimate):** typical sTRX4 boards reserve ~4 lanes for the chipset uplink and commonly wire 1-2 M.2 slots directly to the CPU (4 lanes each) — so realistic GPU-available lanes land around 44-52 of the 64, i.e. **6 GPUs at x8 electrical each (48 lanes) is plausible but board-model-dependent**; x16-each for 6 cards is not possible on 64 lanes regardless of board. x8 electrical is not a meaningful inference-speed penalty for llama.cpp (weights are loaded once; the ongoing per-token traffic across PCIe is small compared to compute), so this is an acceptable tradeoff, not a real bottleneck for this workload. --- ## 6. Power budget vs. the ~200W / one spare 8-pin headroom Official/vendor TDPs: | Card | TDP | Source | |---|---|---| | GTX 1080 (owned) | 180W, one 8-pin | [confirmed 180W, PCIe 3.0 x16, 1× 8-pin](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu) — spec matches NVIDIA's own launch figures reported across multiple outlets incl. Tom's Hardware | | RTX 2080 (owned) | 215W | Cross-checked across gpuzoo/cputronic/notebookcheck spec pages, consistent at 215W | | GT 710 (owned) | ~19W, **no external power connector** (slot power only) | [MSI/EVGA/Zotac GT 710 spec pages](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification) | | RTX 3060 12GB (candidate purchase) | 170W, one 8-pin | [NVIDIA-confirmed 170W TDP, one 8-pin connector](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/) | | RTX 3090 24GB (candidate purchase) | 350W, two 8-pin, [NVIDIA's own RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/) lists 350W and a 750W PSU minimum | NVIDIA official | | R9700 32GB (already deployed / "more of the same") | 300W (per this repo's `server-planing.md`, consistent with AMD's own R9700 product page framing it as a 300W-class card) | in-repo prior research | **Against the stated ~200W / one spare 8-pin budget:** - Adding **one RTX 3060 12GB** (170W, one 8-pin) is the *only* candidate in this list that fits the stated headroom as-is — it uses the one spare connector and stays under 200W. - Adding the already-owned **GTX 1080** (180W) as the classifier's dedicated card also just barely fits (180W ≤ 200W, one 8-pin) — this is a genuinely free option since the card is already owned and its power draw is within budget, unlike every purchase candidate below. - Adding the already-owned **RTX 2080** (215W) **exceeds** the stated 200W headroom by 15W — technically over budget on paper, though real-world draw is usually a bit under rated TDP; flag it as marginal, not safely fitting. - Adding a **second R9700** (300W) or an **RTX 3090** (350W, needs two 8-pin — the user has only one spare) both blow well past the current power budget on both watts and connector count. - **The GT 710 draws no meaningful power (~19W, no PCIe power connector at all)** — it is free from a power-budget standpoint regardless of what else is added. **PSU upgrade trigger:** the very first stage that adds *any* GPU beyond a GTX 1080-class card (180W, one 8-pin) or an RTX 3060 12GB (170W, one 8-pin) exhausts the stated headroom. Any stage that reaches for a second 32GB-class card (R9700 or equivalent) or any 300W+ card **requires a PSU upgrade before that stage**, not after — see the roadmap table in §7 for exactly which stage that is. --- ## 7. Mixed-GPU feasibility: ROCm + CUDA, and is the GT 710 usable at all **ROCm and CUDA are different llama.cpp builds, but that's exactly the pattern already in this repo.** `ghcr.io/ggml-org/llama.cpp` publishes both `server-rocm` and `server-cuda` as separate, independently-built image tags (confirmed present on the [ggml-org container registry](https://github.com/orgs/ggml-org/packages/container/llama.cpp) and documented in [`docs/docker.md`](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md) — *"server-cuda: Same as `server` but compiled with CUDA support"*, *"server-rocm: Same as `server` but compiled with ROCm support"*). You cannot mix backends inside one process/container, but you **can** run one `server-rocm` container pinned to the R9700 and a separate `server-cuda` container pinned to an NVIDIA card, simultaneously, on the same host — this is architecturally identical to today's `llama-server` + `qwen-classifier` two-container split in `docker-compose.yml`, just with a different image tag for the NVIDIA-backed service and NVIDIA's container runtime (`nvidia-container-toolkit` + `--gpus` / device reservation, the CUDA-world equivalent of this repo's `/dev/kfd`+`/dev/dri`+ numeric-GID ROCm pattern documented in [`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md)). None of that doc's ROCm-specific findings (the `GPU_MAX_HW_QUEUES=1` MES firmware workaround, the numeric-GID `group_add` fix) apply to an NVIDIA/CUDA container — those are ROCm-stack-specific bugs, not general multi-GPU-container issues. **Is this an implicit AMD→NVIDIA rebuild, or additive?** Worth surfacing explicitly since the two source plans conflict on this: `server-planing.md` is an AMD-only, ROCm-only 4×R9700 plan. This ticket's plan is **additive/mixed** — keep the R9700 running the main model under ROCm, and bolt on NVIDIA cards under CUDA for secondary roles (classifier, or a second inference GPU for the big model if going the "more of the same type" route means buying NVIDIA instead of more R9700s). Both are internally consistent, but they are different end-states — flag this choice back to the user rather than assuming one. **Splitting the *main* 27B model itself across mixed AMD+NVIDIA silicon in one process is not possible** — llama.cpp's multi-GPU tensor-split only works within a single backend build. To use both an R9700 and an NVIDIA card for the *same* model's layers, all the compute-hosting cards need to be the same backend (all-ROCm or all-CUDA) in that one process. This is why §5's roadmap treats "add GPU capacity to the main model" and "add a GPU for the classifier" as separable purchases with different backend constraints, not a single mixed pool. **Is the GT 710 usable for anything in this pipeline? No.** Reasoning: - 1GB VRAM cannot hold any meaningful fraction of either model's weights (17.6 GB / 2.4-4.3 GB) — even a handful of transformer layers at Q4 quantization exceeds 1GB. - It's Kepler-generation silicon (192 CUDA cores, no tensor cores) — llama.cpp's CUDA backend technically supports pre-Turing cards, but at this VRAM size there's nothing to usefully offload. - It draws power from the PCIe slot only, no external connector — genuinely free to keep installed. - **Plausible actual use: dedicate it as the box's display-output card**, so every compute-capable GPU (R9700, and whichever NVIDIA cards get added) can be fully headless/compute-only with none of their VRAM or a display output tied up driving a monitor — a real, if minor, use for it. This is this doc's own inference from the spec facts above, not a claim found in any primary source. --- ## 8. GPU market pricing vs. the $20-30/GB assumption | Card | VRAM | Backend | Current used-market price (estimate — see caveat) | $/GB | |---|---|---|---|---| | RTX 3060 12GB | 12GB | CUDA | ~$240-300 used (eBay listings, [gpupoet.com tracker](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060): *"from $239"*, [eBay live listings](https://www.ebay.com/shop/rtx-3060-12gb) averaging ~$488 asking but with a $239 floor) | **~$20-25/GB** — matches the stated assumption | | RTX 3090 24GB | 24GB | CUDA | ~$1,010-1,050 used ([bestvaluegpu.com Sep 2026 tracker](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/), [xda-developers coverage](https://www.xda-developers.com/used-rtx-3090-still-best-for-local-ai-in-value/)) | **~$42-44/GB** — well above the stated assumption | | R9700 32GB ("more of the same type") | 32GB | ROCm | **New only — $1,299 MSRP**, street price $1,400-1,585 as of this research ([overclock3d](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/), [pricehistory.app tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm)) — too recent a release (2026) for a used market to exist yet | **~$41-50/GB, and not a used-market price at all** | **Caveat on all three price figures:** these are live marketplace asking-price snapshots pulled via web search on 2026-09-11, not sold-price data or a vendor spec sheet — treat as directional, not exact. eBay asking prices in particular run above realized sale prices. **Correction to the user's stated assumption:** $20-30/GB is a good estimate specifically for **last-generation mainstream used cards** (RTX 3060 12GB fits it almost exactly) but **not** for high-VRAM flagship cards like the RTX 3090 (~1.5-2× that rate) or for "more of the same type" R9700 units, which aren't used-market at all yet and sit even higher per GB than the 3090. If the plan is "cheapest path to more VRAM," multiple RTX 3060 12GB cards (or similar mid-tier used cards) beat one RTX 3090 on $/GB, at the cost of needing more PCIe slots and more total wattage/connectors to reach the same aggregate VRAM — which is exactly the tradeoff the Threadripper lane-count plan in §5 is for. --- ## 9. Step-by-step roadmap All "resulting max context" figures assume `--parallel 2` and the KV precision stated; per §2, the `--ctx-size` value shown is the *total* (pre-division) value to pass to llama-server. | Stage | Hardware change | Est. cost | Backend | Usable VRAM (main-model pool) | Max context @ parallel=2 (`q4_0` KV) | PSU upgrade triggered? | |---|---|---|---|---|---|---| | **0 (current)** | 1× R9700 32GB, in production | $0 | ROCm | 32GB (shared with classifier) | ~131K/slot today at `q8_0` KV (262144 total ÷ 2) | No | | **1 — classifier isolation** | Move classifier onto the already-owned **GTX 1080** (180W, own container, `server-cuda`), freeing the R9700 entirely for the main model. Matches the existing dual-model pattern qwen-code's own docs describe (§10) and this repo's `qwen-classifier` service already implements, just on separate silicon instead of a shared card. | $0 (already owned) | ROCm (main) + CUDA (classifier) | R9700's full 32GB now available to the main model alone | ~262K/slot @ `q8_0` (unchanged ctx-size, no more classifier contention) | **No** — 180W GTX 1080 fits the stated ~200W/one-8-pin headroom | | **2 — second big-model GPU** | Add **one more 32GB-class card** for the main model. Cheapest correct-backend option: a second R9700 (~$1,300-1,585 new, ROCm, same backend as the first — required if tensor-splitting one model across two cards) | ~$1,300-1,585 | ROCm | 64GB pooled | `--ctx-size 500000 --parallel 1` fits at `q4_0` (~33GB) or `q8_0` (~48GB, tight but fits in 64GB) — **not yet 500k×2** | **Yes** — 300W card, no spare 8-pin left after stage 1 | | **3 — reach 500k × 2 / 1M stretch** | No further hardware if stage 2's 64GB pool is used with `--cache-type-k/v q4_0`: `--ctx-size 1000000 --parallel 2` needs ~33-39GB (§4), fits inside 64GB with real headroom for the compute buffer. If `q8_0` KV is required instead (this repo's current quality bar for the main model), the ~51-54GB need is tight-to-marginal on 64GB — a **third** 32GB card (~96GB pool) removes the risk. | $0 (reuses stage 2) or +$1,300-1,585 for a 3rd card if `q8_0` KV is required | ROCm | 64GB (q4_0 case) or 96GB (q8_0 case) | **500k×2 parallel achieved**, and the 1M stretch goal is the *same config* with `--parallel 1` instead of 2 (§2) | Already upgraded at stage 2 | | **4 — optional CPU/lane platform swap** | Only needed if the plan is to keep scaling past 2-3 big cards, or to add several small used cards (RTX 3060 12GB) for extra headroom/throughput rather than raw ctx-size. Swap to **non-PRO Threadripper 3000-series (sTRX4)** — 64 PCIe 4.0 lanes, ~x8-per-slot for up to 6 GPUs (§5). Threadripper PRO 3000WX/5000WX (128 lanes) only if x16-per-card matters or 6+ full-bandwidth slots are wanted. | Used sTRX4 CPU+board: roughly $400-800 combined on the used market (not independently priced in this pass — **estimate**, not cited) | n/a (platform only) | n/a | n/a | Independent of GPU wattage — driven by whatever GPU count/wattage stage 5+ adds | | **5+ — scale-out via small used cards** | Add RTX 3060 12GB units (~$20-25/GB, the assumption that actually holds, §8) instead of more 32GB flagship cards, once lane count (stage 4) supports it — useful for extra parallel slots / throughput beyond the 500k×2 target rather than for raising ctx-size further (500k×2/1M is already met at stage 3). | ~$240-300/card | CUDA (separate container per §6) | +12GB pooled per card, but on a *different backend* from the ROCm main model — usable for extra classifier/small-model capacity or a separate CUDA-backend llama-server instance, not as additional tensor-split VRAM for the ROCm main model | Unchanged for the main model; adds parallel capacity elsewhere | Yes, cumulative — each additional 170W card needs PSU headroom stage 2 already consumed | **Where the existing dual-model pattern sits in this roadmap:** it's stage 1, and it's free. The qwen-code docs pattern (main model + a small, always-resident, non-thinking fast/classifier model — see §10) is already implemented in this repo; the only roadmap-relevant change is *which GPU* the classifier sits on, moving it off the R9700 entirely onto an already-owned NVIDIA card frees the R9700's full 32GB for the 500k×2/1M push instead of splitting it with the classifier as happens today. ### 9.1 Upgrade path, as diagrams Diagram form of the same §9 table and §11.9's dense-vs-Flash-Next call — nothing new is claimed here, this is a visual index back into the cited sections above. **Stage-by-stage hardware path** (PSU-upgrade triggers and target reached called out inline): ```mermaid flowchart TD S0["Stage 0 — today
1x R9700 32GB, ROCm
classifier shares the card
$0"] S1["Stage 1 — classifier isolation
+ GTX 1080 (owned, 180W, CUDA)
R9700 freed for main model
$0 · PSU OK (180W fits ~200W headroom)"] S2["Stage 2 — 2nd big-model GPU
+1x R9700 32GB (ROCm)
64GB pooled
~$1,300-1,585 · PSU UPGRADE REQUIRED (300W, no 8-pin left)"] S3q4["Stage 3a — q4_0 KV
--ctx-size 1,000,000 --parallel 2
~33-39GB, fits in 64GB
$0 (reuses stage 2)"] S3q8["Stage 3b — q8_0 KV (current prod quality)
~51-54GB, tight on 64GB
+1x R9700 -> 96GB removes risk
+~$1,300-1,585"] TARGET(["500k x2 parallel reached
= 1M stretch goal, same VRAM
(--parallel 1 vs 2 is a config flag, §2)"]) S4["Stage 4 — platform swap (optional)
sTRX4 Threadripper 3000, 64 PCIe4 lanes
only needed past 2-3 big cards
~$400-800 (estimate, §9)"] S5["Stage 5+ — scale out
+RTX 3060 12GB cards (CUDA, separate backend)
extra parallel/throughput, not more ctx-size
~$240-300/card · PSU upgrade each card"] S0 --> S1 --> S2 S2 --> S3q4 --> TARGET S2 --> S3q8 --> TARGET TARGET -.->|"only if scaling past this"| S4 --> S5 style TARGET fill:#2e7d32,color:#fff,stroke:#1b5e20 style S2 fill:#8a5a00,color:#fff,stroke:#5c3d00 style S5 fill:#8a5a00,color:#fff,stroke:#5c3d00 ``` **Model choice, and the one open question that could change it** (§11.9): ```mermaid flowchart TD Q{"Goal: 500k-1M ctx
within a $20-30/GB VRAM budget?"} D["Dense Qwen3.8-27B
17.6-54.7GB weights (Q4_K_XL-BF16)
reaches 500k x2 on 2-3 cards,
500k x4 on 2-6 cards depending on quant
(§11.7 tables)"] F{"Try --n-cpu-moe:
offload MoE experts to system RAM?
(untested for this model, §11.8)"} FBAD["Flash-Next, all-GPU weights
111-354GB just for weights
needs 4-13 cards before any KV cost
NOT recommended at this budget (§11.9)"] FGOOD["Flash-Next, experts in system RAM
GPU VRAM could shrink a lot
2.67x cheaper KV/token becomes relevant
UNVERIFIED — prototype on real server first"] CAVEAT["+ real caveat either way:
PR #27742 flags unverified conv branch,
3% QSA divergence, prefill-pos-0-only PLE
(§11.1) — dense model carries no such flag"] Q --> D Q -->|"considering Flash-Next instead"| F F -->|"works well"| FGOOD F -->|"doesn't help / untested"| FBAD FGOOD --> CAVEAT FBAD --> CAVEAT style D fill:#2e7d32,color:#fff,stroke:#1b5e20 style FBAD fill:#8a1c1c,color:#fff,stroke:#5c1212 style FGOOD fill:#8a5a00,color:#fff,stroke:#5c3d00 ``` --- ## 10. Qwen-code's own docs on the fast-model/classifier pattern Fetched directly per the user's link: [qwenlm.github.io/qwen-code-docs/en/users/overview/](https://qwenlm.github.io/qwen-code-docs/en/users/overview/) — **the overview page itself does not describe the dual-model/classifier pattern**; it only covers single-model-provider setup (Alibaba ModelStudio / third-party / custom provider), one model at a time. The actual fast-model/classifier documentation lives on the **Auto Mode** page instead, which this repo's own `fast-model-choice.md` already fetched and cited in detail: [qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/) — summary (see `fast-model-choice.md` §1 for the full quote): a two-stage classifier gate, both stages using "your configured fast model (`/model --fast`)", Stage 1 a ~300ms `{shouldBlock}`-only check, Stage 2 a ~3-5s chain-of-thought review that only runs on a Stage-1 block. Nothing in either page gives a recommended *context size* or *model size* for the fast model beyond what's implied by that latency budget — this repo's own prior research (`fast-model-choice.md`) derived the actual context requirement from qwen-code's source code instead (`packages/core/src/permissions/classifier-transcript.ts`), since the docs pages don't state one. No new information changes that prior doc's conclusion; this section exists to confirm the overview page was checked directly as instructed and doesn't contradict or add to it. --- ## 11. Alternative: Qwen3.8-Flash-Next (MoE, hybrid attention) The user also wants to weigh switching (or adding) **Qwen3.8-Flash-Next** — a 125B-total/6B-active MoE with a hybrid recurrent-attention architecture — against staying on dense Qwen3.8-27B, and separately wants this section to cover **going up in weight quant** (Q4_K_XL → Q6_K_XL → Q8_0 → BF16/fp16) for *both* models, not just Q4. Feasibility first, since it gates everything else. ### 11.1 Feasibility verdict: supported, but immature — read before trusting any number below Checked directly against the primary sources the task named: - **llama.cpp mainline support exists.** [PR #27742](https://github.com/ggml-org/llama.cpp/pull/27742) ("model: add Qwen3.8-Flash-Next (qwen4exp)") was **merged into `master` on 2026-08-27** by ngxson. It adds the full architecture: Gated DeltaNet layers (sigmoid-gated linear attention), QSA ("Qwen Sparse Attention", operating at micro-block granularity), hyper-connections, and the PLE n-gram embedding table. `llama.cpp`'s own docs list CPU/CUDA/Metal/ROCm as supported backends for it — this is not a CUDA-only feature. - **This repo's pinned image is a floating tag, not a version pin.** `docker-compose.yml` runs `ghcr.io/ggml-org/llama.cpp:server-rocm` with no date/digest suffix — a rolling "latest ROCm server build" tag, not a release version. The merge is from 2026-08-27, and today is 2026-09-11 (~2 weeks later), so a **fresh pull** of `server-rocm` should include it — but whatever image is already cached/running on the R9700 box may predate the merge. **Action before touching this model on the server: `docker compose pull llama-server` and check the startup log's build/commit banner is dated on/after 2026-08-27**, not just "the tag says server-rocm." - **Real, primary-source-flagged immaturity — this is the part that should temper enthusiasm.** The PR's own description/review discussion states: *"The conv branch itself is still numerically unverified because the fixture zeroes its weights"*; QSA sparse attention *"diverges on 3 percent of positions"* above its budget threshold; the PLE depthwise convolution *"is exact only for a prefill that starts at position 0"* (i.e. correctness is not guaranteed once `--cache-reuse`/prompt-caching is in play — a flag this repo already turns on for the dense model per the latest commit). None of that is disqualifying, but it is a primary-source admission that this is a fresh, not-fully-verified implementation, not a mature, widely-battle-tested one like the dense Qwen3.8-27B path. - **Multi-slot serving needs an explicit new flag.** The same PR states: *"`set_input_qsa` asserted `n_stream == 1`, so llama-server could not serve this model with more than one slot unless `-kvu` was passed."* Per the server README (§2), `--kv-unified`/`-kvu` defaults to enabled **only when slot count is auto** (`-1`). This repo's compose file sets `--parallel ${LLAMA_PARALLEL:-2}` **explicitly** (not auto) — so adopting Flash-Next with `--parallel` > 1 requires **adding `--kv-unified` (or `-kvu`) to the launch flags**, a real deploy-time change, not something that "just works" by copying today's flag set onto a new model file. **Verdict: yes, runnable** on this repo's backend (ROCm, mainline, no dev branch needed) as long as the image is pulled after 2026-08-27 and `-kvu` is added for multi-slot use — but treat it as **usable-with-caution**, not a drop-in swap, given the PR author's own unresolved-correctness notes. ### 11.2 Architecture, verified against `config.json` directly Fetched from `Qwen/Qwen3.8-Flash-Next`'s `config.json` (unsloth's GGUF repo repackages the same base model): `num_hidden_layers=48`, `hidden_size=2560`, `num_attention_heads=24`, `num_key_value_heads=2`, `head_dim=256`, `max_position_embeddings=262144` (same native/extensible-to-1M framing as the dense model — same YaRN quality caveat from §3 applies here too, unverified past native length), `num_experts=512`, `num_experts_per_tok=10`, and the linear-attention head config: `linear_num_key_heads=16`, `linear_num_value_heads=48`, `linear_key_head_dim=128`, `linear_value_head_dim=128`. Layer pattern (confirmed both from the model card's own description and `config.json`'s `full_attention_interval=4`): every 4th layer is full/QSA attention, the other 3 are Gated DeltaNet — **12 of 48 layers grow a real KV cache; the other 36 have a fixed-size recurrent state that does not grow with context length.** (24% full-attention layers vs. the dense model's 16-of-64 = 25% — similar ratio, but the *absolute* per-layer KV cost differs because `num_key_value_heads` is 2 here vs. 4 on the dense model — see below.) ### 11.3 Per-token growing-KV-cache cost ``` 12 full-attention layers × 2 (K+V) × 2 kv_heads × 256 head_dim × 2 bytes (fp16) = 24 KiB/token ``` | Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 | |---|---|---|---| | 262,144 (native) | ~6.0 GiB | ~3.0 GiB | ~1.5 GiB | | 500,000 | ~11.4 GiB | ~5.7 GiB | ~2.9 GiB | | **1,000,000 (500k×2, or 1M stretch)** | **~22.9 GiB** | **~11.4 GiB** | **~5.7 GiB** | | **2,000,000 (500k×4)** | **~45.8 GiB** | **~22.9 GiB** | **~11.4 GiB** | `--cache-type-k/v` are the same generic llama.cpp KV-cache-quantization flags used elsewhere in this doc; nothing in the PR or the server README suggests they're handled differently for the 12 full-attention layers of a hybrid model — they quantize the same growing K/V buffers as on a plain transformer. (No primary source explicitly confirms this for *this* architecture specifically — flagged as a reasonable extrapolation, not a directly-cited fact, same caveat class as this doc's other estimates.) ### 11.4 Fixed (non-growing) recurrent state — Gated DeltaNet layers The 36 Gated DeltaNet layers each keep a fixed-size recurrent state (an outer-product-style key×value matrix per head) that does **not** scale with context length — only with slot/sequence count. Sized from `config.json`'s linear-attention head params: ``` 36 layers × linear_num_value_heads(48) × linear_key_head_dim(128) × linear_value_head_dim(128) × 4 bytes (fp32 state) ≈ 36 × 48 × 128 × 128 × 4 bytes ≈ 108 MiB per slot ``` This is **this doc's own derivation from the published head-dimension params, not a value pulled directly from llama.cpp source or docs** — the PR text confirms the state exists per-stream/per-slot but doesn't publish an exact byte formula, so treat the ~108 MiB/slot figure as an estimate, medium confidence. Even at 4 parallel slots that's under half a gigabyte — **negligible** next to both the growing KV cache (GBs) and the weights (tens to hundreds of GB) computed below. The headline implication holds regardless of the exact multiplier: Flash-Next's "big memory line item" is the MoE weights, not the attention state of any kind. ### 11.5 Magnitude vs. the dense model — how much cheaper is KV, really At the same total ctx-size, Flash-Next's growing KV cache is **24 KiB/token vs. the dense model's 64 KiB/token — 2.67× smaller**, i.e. Flash-Next's KV budget is **37.5%** of the dense model's at identical context length. This is a real, significant win *for the KV-cache line item specifically* — but see §11.9: it's a much smaller slice of a much bigger total, because the weights move the other way by a far larger factor. ### 11.6 Weight sizes — verified from each unsloth GGUF repo's actual file listing Fetched directly from the HF file trees (not estimated from ratios), current as of this research pass: | Quant tier | Qwen3.8-27B (dense) | Qwen3.8-Flash-Next (MoE) | |---|---|---| | Q4_K_XL (`UD-Q4_K_XL`) | **17.6 GB** (existing baseline) | **111.4 GB** (4 parts: 10.9MB + 49.9GB + 49.4GB + 12.1GB) | | Q6_K_XL (`UD-Q6_K_XL`) | **25.3 GB** | **169 GB** (6 parts) | | Q8_0 | **29 GB** | **188 GB** (6 parts) | | BF16/fp16 | **54.67 GB** (50GB + 4.67GB, 2 parts) | **354 GB** (8 parts) | Sources: [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main) and its `BF16/` subfolder; [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main) and its `UD-Q4_K_XL/`, `UD-Q6_K_XL/`, `Q8_0/`, `BF16/` subfolders (per-file sizes summed). The preliminary Q8_0 figure floated before this research pass (~192GB) was slightly high — the real listing sums to **188 GB**; everything else in the preliminary list was accurate to within rounding. **The weight-quant axis and the KV-cache-quant axis are independent knobs.** Raising weight quality (Q4_K_XL → BF16) does not require raising `--cache-type-k/v` — the two flags are unrelated, and this repo already proves that pattern works (`q8_0` KV cache is deployed today against `Q4_K_XL` weights). A user chasing **maximum output quality** can run e.g. **BF16 weights + `q4_0` KV cache** — full-precision weights for quality, still-compressed KV for context budget — or any other combination in the tables below; nothing about picking a higher weight quant forces a matching KV precision. ### 11.7 Total VRAM: does it fit, across quant tiers and both parallelism targets All totals = weights + growing KV cache + an estimated **+3-6 GB** compute-buffer/runtime overhead (same estimate band as §4, carried over — not re-derived for this architecture; flagged medium confidence there too). "Cards" = ceil(total ÷ 32GB), i.e. how many R9700-class 32GB cards it takes. #### Dense Qwen3.8-27B — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`) | Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) | |---|---|---|---| | Q4_K_XL (17.6GB) | ~79-85GB (**3**) | ~48-54GB (**2**) | ~33-39GB (**2**) | | Q6_K_XL (25.3GB) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) | ~44-47GB (**2**) | | Q8_0 (29GB) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) | ~47-50GB (**2**) | | BF16 (54.67GB) | ~119-122GB (**4**) | ~88-91GB (**3**) | ~73-76GB (**3**) | #### Dense Qwen3.8-27B — 500k × 4 parallel (`--ctx-size 2,000,000`) | Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) | |---|---|---|---| | Q4_K_XL (17.6GB) | ~143-146GB (**5**) | ~82-85GB (**3**) | ~51-54GB (**2**) | | Q6_K_XL (25.3GB) | ~150-153GB (**5**) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) | | Q8_0 (29GB) | ~154-157GB (**5**) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) | | BF16 (54.67GB) | ~180-183GB (**6**) | ~119-122GB (**4**) | ~88-91GB (**3**) | #### Qwen3.8-Flash-Next — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`) | Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) | |---|---|---|---| | UD-Q4_K_XL (111.4GB) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) | ~120-123GB (**4**) | | UD-Q6_K_XL (169GB) | ~195-198GB (**7**) | ~183-186GB (**6**) | ~178-181GB (**6**) | | Q8_0 (188GB) | ~214-217GB (**7**) | ~202-205GB (**7**) | ~197-200GB (**7**) | | BF16 (354GB) | ~357-360GB (**12**) | ~357-360GB (**12**) | ~357-360GB (**12**) | #### Qwen3.8-Flash-Next — 500k × 4 parallel (`--ctx-size 2,000,000`) | Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) | |---|---|---|---| | UD-Q4_K_XL (111.4GB) | ~160-163GB (**6**) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) | | UD-Q6_K_XL (169GB) | ~218-221GB (**7**) | ~195-198GB (**7**) | ~183-186GB (**6**) | | Q8_0 (188GB) | ~234-237GB (**8**) | ~211-214GB (**7**) | ~199-202GB (**7**) | | BF16 (354GB) | ~397-400GB (**13**) | ~377-380GB (**12**) | ~366-369GB (**12**) | (Flash-Next's KV precision barely moves the total at any weight quant above `UD-Q6_K_XL` — the weights so dominate the budget that KV quantization stops mattering for the "how many cards" question. This is the clearest signal in this whole section: for Flash-Next, the weight-quant choice is the entire hardware-sizing decision; for the dense model, KV precision still matters a lot.) ### 11.8 CPU MoE-expert offload — the one lever that could change this calculus Flash-Next is a 512-expert/10-active-per-token MoE, and llama.cpp has a purpose-built flag for exactly this shape of model, confirmed directly from the server README: **`--n-cpu-moe`** — *"keep the Mixture of Experts (MoE) weights of the first N layers in the CPU"* — plus the more general **`--override-tensor`** (*"override tensor buffer type"*, pattern-matched by tensor name) that the same flag is built on top of. Both are generic, architecture-agnostic llama.cpp mechanisms (they match on tensor name patterns, not model type), so there's no reason to expect them not to apply to Flash-Next's MoE tensors specifically — but this pass found **no primary source that has actually tested `--n-cpu-moe` against this specific qwen4exp architecture**, so treat "it works here" as plausible, not confirmed. If it does work as expected, this changes the whole weight-VRAM picture in §11.7: the ~90-95% of Flash-Next's weight footprint that's MoE expert tensors could live in system RAM while attention projections, the shared/non-expert tensors, and the full KV cache stay on GPU — meaning a much smaller GPU-VRAM number than the "all weights on GPU" tables above, at the cost of PCIe/RAM-bandwidth-bound inference speed for whichever experts get selected per token (this repo has no benchmark of that tradeoff, and it's highly system-RAM-bandwidth-dependent, so no number is given here — flagged as an escape hatch worth prototyping directly on the server, not something this research values responsibly without a real test run). ### 11.9 Net recommendation: dense Qwen3.8-27B vs. Flash-Next, for this user's stated goal **Net loss for this user's goal, as things stand — stay on dense Qwen3.8-27B.** Reasoning: - The user's target (500k×2 or 500k×4, on a $20-30/GB-VRAM budget, GPUs in 32GB increments) is a **VRAM-budget-constrained** goal, and §11.7 shows Flash-Next's *weights alone* (111-354GB depending on quant) dwarf the entire dense-model total-VRAM figure from §4/§11.7 (33-183GB depending on quant) at every parallelism target. Flash-Next's much cheaper per-token KV cache (§11.5, real and verified) is a rounding error next to that weight-size gap — the "2.67× cheaper KV" win doesn't come close to offsetting a "6-20× larger weight footprint," so at $20-30/GB-VRAM the *dense* model reaches 500k×2 or 500k×4 for a fraction of the card count and dollar cost that Flash-Next needs even at its lowest usable quant (`UD-Q4_K_XL`, 4-5 cards minimum) — before even factoring in §11.1's immaturity flags. - The one scenario that could flip this verdict is `--n-cpu-moe` actually working well for this architecture (§11.8) — if most of those 111-354GB of expert weights can sit in system RAM at acceptable throughput, Flash-Next's GPU-VRAM number could shrink dramatically and its real KV-cache advantage would start to matter. That is untested here and shouldn't be assumed; it's the one concrete next step worth trying on the actual server before ruling Flash-Next out permanently. - Independent of VRAM: §11.1's primary-source-flagged correctness caveats (unverified conv branch, 3%-divergence QSA, prefill-position-0-only PLE exactness) are a real quality/stability risk on a production coding-agent stack that dense Qwen3.8-27B simply doesn't carry, since it's been running in this repo already. --- ## 12. 4-parallel × 500k scenario — all four combinations side by side Per §2's already-established, cited rule (`n_ctx_per_seq = n_ctx / n_parallel`, [ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681)), the same division applies at 4 slots: **500k tokens on each of 4 parallel slots requires `--ctx-size 2,000,000`** — double the 2-parallel target's `--ctx-size 1,000,000`, for the same reason 500k×2 needed double 262,144. This isn't a new mechanism, just the same formula at `--parallel 4`. Full per-quant-tier tables for all four combinations are in §11.7 above (dense×2, dense×4, Flash-Next×2, Flash-Next×4 are each their own table there). Headline comparison at the KV precision already proven in production in this repo (`q8_0`) and each model's respective current/cheapest-usable weight quant: | Scenario | `--ctx-size` | Weight quant | q8_0-KV total VRAM | Cards (32GB) | |---|---|---|---|---| | Dense × 2 (or 1M stretch) | 1,000,000 | Q4_K_XL (17.6GB, current) | ~48-54GB | **2** | | Dense × 4 | 2,000,000 | Q4_K_XL (17.6GB, current) | ~82-85GB | **3** | | Flash-Next × 2 (or 1M stretch) | 1,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~126-129GB | **4**, edge | | Flash-Next × 4 | 2,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~137-140GB | **5** | **4-parallel × 500k reachability against this repo's existing roadmap stages (§9):** - **(a) Current 1×R9700 32GB:** none of the four combinations fit — not even dense×2 at any weight/KV quant (§4's own conclusion, unchanged). - **(b) The 2-3×R9700 roadmap already proposed in §9 (64-96GB):** covers **dense×2 fully** (stage 3, as already established) and **dense×4 at `q4_0` KV with Q4_K_XL or Q6_K_XL weights** (~51-62GB, fits in 64-96GB) — but **not** dense×4 at higher weight quants (Q8_0/BF16 need 3-6 cards depending on KV precision, per §11.7's dense×4 table) and **not any Flash-Next scenario** (minimum is 4 cards/128GB even at the cheapest usable quant and tightest KV). - **(c) The full 4-6×R9700 stretch scenario** (`server-planing.md`'s original plan, 128-192GB pooled): covers **dense×4 at every weight quant up to BF16** (worst case ~91GB at BF16+q4_0, well inside 128GB) and **Flash-Next×2 at `UD-Q4_K_XL`** (126-140GB, fits a 5-card/160GB build, tight on a 4-card/ 128GB one) — but **not** Flash-Next×4 at any weight quant above `UD-Q4_K_XL`, and not Flash-Next at `BF16` under any parallelism (needs 12-13 cards, an entirely different scale of build than anything in this doc's roadmap). --- ## Sources - [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) — `--ctx-size`, `--parallel`, `--cache-type-k/v`, `--kv-unified`, `--cache-reuse`, `--n-cpu-moe`, `--override-tensor` flag definitions - [ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681) — real server log proving `n_ctx_per_seq = n_ctx / n_parallel` - [ggml-org/llama.cpp#27742](https://github.com/ggml-org/llama.cpp/pull/27742) — "model: add Qwen3.8-Flash-Next (qwen4exp)", merged 2026-08-27; architecture details, `n_stream == 1` / `-kvu` multi-slot requirement, and the conv-branch/QSA-divergence/PLE-prefill correctness caveats - [Qwen/Qwen3.8-27B config.json](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json) - [Qwen/Qwen3.8-Flash-Next config.json](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/raw/main/config.json) - [Qwen/Qwen3-4B-Instruct-2507 config.json](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json) - [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main) — 17.6GB (Q4_K_XL), 25.3GB (Q6_K_XL), 29GB (Q8_0), 54.67GB (BF16) weight sizes - [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main) — 111.4GB (UD-Q4_K_XL), 169GB (UD-Q6_K_XL), 188GB (Q8_0), 354GB (BF16) weight sizes, summed from each quant's per-file listing - [ggml-org/llama.cpp docs/docker.md](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md) — `server-cuda`/`server-rocm` separate image tags - [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html) - [Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2) - [PCWorld — Threadripper PRO launch, 128 PCIe lanes](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html) - [NVIDIA — GeForce RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/) - [Lowyat.net — RTX 3060 official 170W TDP](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/) - [BuildMyServer — GTX 1080 180W/PCIe3.0/1×8-pin spec listing](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu) - [MSI — GT 710 1GD5 LP spec page](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification) - [overclock3d — AMD Radeon AI PRO R9700 $1,299 MSRP](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/) - [pricehistory.app — R9700 street price tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm) - [bestvaluegpu.com — RTX 3090 used price tracker, Sep 2026](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/) - [gpupoet.com — RTX 3060 12GB used listings](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060) - [Qwen Code docs — overview](https://qwenlm.github.io/qwen-code-docs/en/users/overview/) - [Qwen Code docs — Auto Mode (fast-model pattern)](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/) - This repo: [`docker-compose.yml`](../../docker-compose.yml), [`.env.example`](../../.env.example), [`docs/server-planing.md`](../server-planing.md), [`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), [`qwen3.8-27b-tool-calling.md`](qwen3.8-27b-tool-calling.md), [`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md), [`fast-model-choice.md`](fast-model-choice.md) ## Confidence/uncertainty summary - **High confidence:** the KV-cache-per-token formulas for both dense models and Flash-Next (computed directly from each model's own `config.json`, same method this repo's prior research already used and cross-checked); the `n_ctx_per_seq = n_ctx / n_parallel` division behavior (directly evidenced by a real server log in a llama.cpp GitHub issue, and independently already documented in this repo's own `.env.example`) — and confirmed to apply identically at `--parallel 4` since the mechanism is parallel-count-agnostic; official TDP figures for GTX 1080, RTX 2080, RTX 3060, RTX 3090, GT 710 (each cross-checked against 2+ independent spec listings or the vendor's own product page); the TR4-is-PCIe3/sTRX4-is-PCIe4 generational split (direct launch-coverage quote); the existence of separate `server-cuda`/`server-rocm` llama.cpp image tags; Qwen3.8-Flash-Next's `config.json` architecture params and PR #27742's merge date/status and its own stated correctness caveats and `-kvu` multi-slot requirement (all directly quoted from the primary source); the weight file sizes for both models at all four quant tiers (summed directly from each HF repo's real file listing, not estimated). - **Medium confidence:** the compute-buffer/runtime-overhead estimate in §4/§11.7 (+3-6 GiB) — extrapolated from one in-repo incident's before/after numbers, not a llama.cpp-documented formula, and carried over to Flash-Next without re-derivation for its different architecture; the Gated DeltaNet fixed recurrent-state size in §11.4 (~108 MiB/slot) — this doc's own derivation from the published head-dimension config, not a value found in llama.cpp source or docs; whether `--cache-type-k/v` quantization applies identically to Flash-Next's 12 full-attention layers as it does to a plain transformer (reasonable extrapolation, not directly confirmed for this architecture); whether `--n-cpu-moe`/`--override-tensor` actually work against Flash-Next's specific MoE tensor layout (architecture-agnostic mechanism, but untested against this model by any primary source found); real-world PCIe lane availability for 6 GPUs on a specific sTRX4 board (§5) — no single board's exact lane map was fetched, this is a reasonable-but-unverified estimate from typical sTRX4 board behavior. - **Low confidence / explicitly estimated, not cited fact:** all used-GPU marketplace pricing (§8) — live asking-price snapshots from a single search pass, not sold-price data; the used sTRX4 CPU+motherboard combo price in the roadmap's stage 4 (§9) — not researched at all in this pass, flagged as a placeholder estimate; whether YaRN-scaled 500k/1M context actually holds output quality for either Qwen3.8-27B or Qwen3.8-Flash-Next — no primary source (Qwen's own docs included) publishes long-context quality benchmarks past the 262,144 native length for either model, so this is a known-unknown carried forward from each model card's "YaRN-extensible" claim, not a verified capability; whether the specific `ghcr.io/ggml-org/llama.cpp:server-rocm` image currently cached on this repo's server actually postdates PR #27742's 2026-08-27 merge — not checked against the live server in this pass, flagged as an action item in §11.1 rather than a confirmed fact.