From 2250e804db125b9edc123e8ea216ff2a12e00ceb Mon Sep 17 00:00:00 2001 From: Haylan Date: Fri, 11 Sep 2026 13:15:30 +0200 Subject: [PATCH] Implement code changes to enhance functionality and improve performance --- ...k-context-2-parallel-agents-gpu-roadmap.md | 737 ++++++++++++++++++ 1 file changed, 737 insertions(+) create mode 100644 docs/research/500k-context-2-parallel-agents-gpu-roadmap.md diff --git a/docs/research/500k-context-2-parallel-agents-gpu-roadmap.md b/docs/research/500k-context-2-parallel-agents-gpu-roadmap.md new file mode 100644 index 0000000..8a2ce39 --- /dev/null +++ b/docs/research/500k-context-2-parallel-agents-gpu-roadmap.md @@ -0,0 +1,737 @@ +# Research: hardware roadmap to 500k-token context × 2 parallel agents (1M stretch) + +**Date:** 2026-09-11 + +**Question:** What VRAM does 500k-token context × 2 parallel llama-server slots (and a 1M-token +stretch goal) actually cost for the Qwen3 family, and what hardware roadmap gets there from the +current single-R9700 setup — given the user's stated plan to add an older (PCIe 4.0) Threadripper +for lane count, reuse existing RAM/PSU (~200W headroom / one spare 8-pin), mix in already-owned +NVIDIA cards (GTX 1080 8GB, RTX 2080 8GB, GT 710 1GB) for the classifier role, and price used GPUs +at roughly $20-30/GB VRAM? + +**Answer, short version:** The two goals ("500k × 2 parallel" and "1M stretch") turn out to need +**the same total VRAM budget** — because of how llama-server's `--ctx-size` and `--parallel` interact +(§2), 500k × 2 slots and a single 1M-token slot both require setting `--ctx-size 1000000`. At +`q4_0`-quantized KV cache that's **~33 GB** (weights + KV) for Qwen3.8-27B, at `q8_0` it's **~48 GB**, +at fp16 it's **~79 GB** — before compute-buffer overhead. That does not fit on the current single +32GB R9700 at any KV precision, and comfortably fits on two 32GB-class cards only at `q8_0`/`q4_0`. +The user's $20-30/GB pricing intuition holds for last-gen used consumer cards (RTX 3060 12GB) but +**not** for RTX 3090 24GB (~$44/GB currently) or a second R9700 (~$41/GB, new — no used market yet +for a card released mid-2026). The stated Threadripper plan needs to specifically target the +**non-PRO Threadripper 3000 series on sTRX4** (64 lanes, PCIe 4.0) — older Threadripper on the +original TR4 socket (1000/2000 series) is PCIe 3.0 only, which doesn't match the user's own PCIe 4.0 +requirement. The power budget (~200W / one spare 8-pin) is exhausted by a *single* mid-tier used GPU +addition — a PSU upgrade is not optional past the very first stage. See §7 for the roadmap. + +--- + +## 1. Current state (from this repo) + +From `docker-compose.yml` and `.env.example` at the repo root: + +- **Main model:** `Qwen3.8-27B-UD-Q4_K_XL.gguf` (17.6 GB weights), `--ctx-size 262144`, + `--parallel 2`, `--flash-attn on`, `--cache-type-k q8_0 --cache-type-v q8_0`, `--n-gpu-layers 999`, + on one AMD Radeon AI PRO R9700 (32GB, ROCm/HIP, `gfx1201`). +- **Classifier ("fast") model:** `Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`, `--ctx-size 65536`, + `--parallel 1`, `--n-gpu-layers 28` (partial offload), `--cache-type-k/v q4_0`, its own container on + the *same* R9700, sharing VRAM with the main model — see + [`docker-compose.yml`](../../docker-compose.yml) lines ~58-99 and + [`fast-model-choice.md`](fast-model-choice.md). +- `.env.example` already documents the exact fact this research turns on: *"Each slot gets + `LLAMA_CTX_SIZE / LLAMA_PARALLEL` tokens of context"* — i.e. today's 262144 ctx-size ÷ 2 parallel + slots means each real request only gets **~131K tokens**, not the full 262144, confirmed in-repo + before any external source was checked. +- Prior research already worked out the KV-cache formula for this exact model + ([`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md)) — this doc reuses and extends that math for the + 500k/1M targets rather than re-deriving it. + +`docs/server-planing.md` describes a **different, earlier plan**: a 4× AMD Radeon AI PRO R9700 rig +on a Gigabyte MZ32-AR0 (single-socket SP3/EPYC, 128 PCIe 4.0 lanes), fully AMD/ROCm. The user's plan +in this ticket is not that — it pivots toward an older **Threadripper** (SP3's sibling desktop-HEDT +socket family, not SP3 itself) and explicitly wants to mix in already-owned **NVIDIA** cards. These +two plans are **not the same build** and, per §6, ROCm and CUDA cards cannot share one llama.cpp +process — they can only coexist as separate containers on separate cards. Treat `server-planing.md` +as superseded context, not the active plan, unless the user says otherwise. + +--- + +## 2. llama-server parallelism: does each slot get its own full `--ctx-size`, or is it divided? + +**Divided.** This is the single fact that changes the whole budget by 2×, confirmed from three +independent primary sources: + +1. **This repo's own `.env.example`** (quoted above) already documents it for the current deployment. +2. **llama.cpp's own server README**, fetched directly + ([`tools/server/README.md`](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)): + `--ctx-size (-c)`: *"size of the prompt context (default: 0, 0 = loaded from model)"*; + `--parallel (-np)`: *"number of server slots (default: -1, -1 = auto)"* — the docs list these as + independent flags, but don't spell out the division themselves. +3. **A real user's server log**, quoted verbatim in + [ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681), is the actual proof: + running with `--ctx-size 327680 --parallel 6` produces `n_ctx = 327680`, + `n_ctx_per_seq = 54613` — i.e. `327680 / 6 ≈ 54613`. The reporter explicitly asked for a + `--ctx-size-per-seq`-style flag to *avoid* this division; no such flag exists as of the fetch date. + +Practical consequence: **to get 500,000 usable tokens on each of 2 parallel slots, `--ctx-size` must +be set to 1,000,000, not 500,000.** The KV cache is sized off the *total* `--ctx-size` +(`--kv-unified`, on by default when slots are auto per the README's `-kvu` entry, uses one shared +pool sized to the full `n_ctx`) — so the VRAM cost of "500k × 2 parallel" and "one 1M-token slot" +is **identical**: both require `--ctx-size 1000000`. This is a genuinely useful finding for the +roadmap — reaching the 500k×2 target and the 1M stretch goal cost the same VRAM; the only difference +is `--parallel 1` vs `--parallel 2` at deploy time, a config change with zero extra hardware cost. + +`--cache-type-k` / `--cache-type-v` accept `f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1` +(default `f16`), per the same README fetch. The repo already uses `q8_0` on the main model and `q4_0` +on the classifier, so both quantization tiers used in the math below are already-proven-working +configurations in this stack, not hypothetical flags. + +--- + +## 3. KV-cache math per model + +### Qwen3.8-27B (hybrid Gated-DeltaNet / attention) + +Reusing the architecture params already pulled from +[`Qwen/Qwen3.8-27B/config.json`](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json) in +[`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), re-verified directly for this doc: `num_hidden_layers=64`, +`full_attention_interval=4` → **16 of 64 layers are standard KV-caching attention**, the other 48 are +Gated DeltaNet linear-attention layers with a small, context-length-*independent* recurrent state +(tens of MB total, negligible next to the attention KV cache — ignored below). +`num_key_value_heads=4` (GQA), `head_dim=256`. Native context `max_position_embeddings=262144` +(YaRN-extensible to 1M per the model card — **both the 500k and 1M targets exceed native context and +require RoPE/YaRN scaling**, which is a real quality caveat, not just a memory one — Qwen has not +published independent long-context quality benchmarks past native length that this research found). + +Per-token KV cache, fp16, both K and V, across the 16 full-attention layers: + +``` +16 layers × 2 (K+V) × 4 kv_heads × 256 head_dim × 2 bytes = 64 KiB/token +``` + +| Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 | +|---|---|---|---| +| 262,144 (current) | ~16.0 GiB | ~8.0 GiB | ~4.0 GiB | +| 500,000 | ~30.5 GiB | ~15.3 GiB | ~7.6 GiB | +| **1,000,000 (500k×2, or 1M stretch)** | **~61.0 GiB** | **~30.5 GiB** | **~15.3 GiB** | + +(`q8_0` is 8-bit vs. fp16's 16-bit → exactly half; `q4_0` is 4-bit → exactly quarter, per llama.cpp's +own cache-type byte widths.) + +### Qwen3-4B-Instruct-2507 (plain GQA transformer, classifier role) + +From [`Qwen/Qwen3-4B-Instruct-2507/config.json`](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json) +(already pulled in [`fast-model-choice.md`](fast-model-choice.md)): `num_hidden_layers=36` — every +layer is standard attention here (no hybrid split), `num_key_value_heads=8`, `head_dim=128`. + +``` +36 layers × 2 (K+V) × 8 kv_heads × 128 head_dim × 2 bytes = 144 KiB/token +``` + +The classifier's real transcript ceiling is ~40-50K tokens (qwen-code's own +`MAX_TRANSCRIPT_MESSAGES=40` × `MAX_HISTORICAL_ACTION_CHARS=4000`, per `fast-model-choice.md` §"what +actually shipped") — nowhere near 500k/1M, so the classifier does **not** need to grow for this +roadmap; it stays exactly as deployed today, on its own small allocation. Per-token cost is included +here only because it feeds the "does the classifier's dedicated GPU need to change" question in §6. + +--- + +## 4. Total VRAM budget: 500k × 2 parallel, and the 1M stretch + +Weights: `Qwen3.8-27B-UD-Q4_K_XL.gguf` is **17.6 GB**, confirmed directly from the +[unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main) +(already verified in `qwen3.8-27b-quant.md`). + +Per §2, both "500k × 2 parallel" and "1M stretch" require `--ctx-size 1000000` — same KV budget: + +| KV precision | KV cache | + weights (17.6 GB) | + est. compute-buffer/runtime overhead* | **Realistic total** | +|---|---|---|---|---| +| fp16 (default) | 61.0 GiB | 78.6 GiB | +3-6 GiB | **~82-85 GB** | +| q8_0 (proven in this stack today) | 30.5 GiB | 48.1 GiB | +3-6 GiB | **~51-54 GB** | +| q4_0 (proven in this stack today, on the classifier) | 15.3 GiB | 32.9 GiB | +3-6 GiB | **~36-39 GB** | + +\* *Estimate, not a cited figure* — llama.cpp's flash-attention compute buffer scales closer to +linear than the unfused-attention path, per this repo's own measured note in `docker-compose.yml`'s +`qwen-classifier` comment (unfused attention buffers ballooned unexpectedly at 65536 ctx; flash-attn +fixed it). `--flash-attn on` is already the deployed default for the main model, so the linear-ish +regime applies, but no primary source gives an exact formula for this buffer size at 1M context — the ++3-6 GiB band is this doc's estimate based on the ratio observed in that in-repo incident, not a +llama.cpp-documented number. Budget for the high end of that range when sizing hardware. + +**Bottom line:** at `q4_0` KV (the most aggressive, already-proven-in-this-repo tier), 500k×2 / +1M needs **~36-39 GB** total VRAM for the 27B model alone. That does not fit one 32GB card at any +precision — it needs at least two 32GB-class cards, or one ≥40GB card. At `q8_0` (the precision this +repo already runs in production for quality reasons), budget **~51-54 GB** — two 32GB cards (64GB +pooled) clears this with room to spare; a single 48GB-class card would not. + +(§11 below extends this table to higher weight-quant tiers — Q6_K_XL, Q8_0, BF16 — for users who want +better output quality than `Q4_K_XL`, and to a Flash-Next alternative architecture; see §11.6-§11.7.) + +--- + +## 5. CPU/motherboard: which Threadripper generations give PCIe 4.0, and how many lanes for GPUs + +AMD's own product/chipset pages, cross-checked against the launch reviews that quote them directly: + +| Platform | Socket | PCIe generation | Total CPU-provided lanes | +|---|---|---|---| +| Threadripper 1000/2000 series ("1920X", "2950X", etc.) | **TR4** | **PCIe 3.0 only** | 60-64 | +| Threadripper 3000 series (3960X/3970X/3990X) | **sTRX4** | **PCIe 4.0** | 64 | +| Threadripper 7000 series (non-PRO) | sTR5 | PCIe 5.0 (48 lanes) + PCIe 4.0 (24-32 lanes) | ~72-80 | +| Threadripper PRO 3000WX/5000WX | sWRX8 | PCIe 4.0 | **128** | +| Threadripper PRO 7000WX | sTR5 (WRX90) | PCIe 5.0 (128 lanes) + a few PCIe 3.0 | **128** | + +Sources: [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html) +(sWRX8/socket listing), corroborated by +[Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch coverage](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2) +(*"the 3rd Gen TR CPUs carry the same 64 PCIe lanes but double bandwidth by moving from Gen 3.0 to +Gen 4.0"* — explicit confirmation TR4/1000-2000-series is PCIe 3.0 while sTRX4/3000-series is PCIe +4.0), [PCWorld — Threadripper PRO launch](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html) +(*"128 PCIe lanes"* for PRO). + +**This directly matters for the user's plan.** "An older Threadripper... for more PCIe lanes" is +ambiguous between two real, very different chips: + +- **TR4 (1000/2000 series)** — cheapest used option, but **PCIe 3.0** — does not meet the user's own + stated PCIe 4.0 requirement, and PCIe 3.0 x8 per GPU roughly halves inter-GPU/host transfer + bandwidth (matters more for training/tensor-parallel than for llama.cpp's inference-time layer + splitting, but still a real downgrade vs. the R9700's native PCIe 5.0). +- **sTRX4 (3000 series, non-PRO)** — the correct "older Threadripper with PCIe 4.0" target: 64 lanes, + 4-5 years old, real used-market availability, no PRO price premium. +- **Threadripper PRO (3000WX/5000WX)** doubles the lane count to 128 but at meaningfully higher used + cost (workstation-tier, lower volume, sWRX8 boards are pricier than sTRX4/TRX40 boards) — worth it + only if 6 full-bandwidth (x16) GPU slots are actually needed; at x8-per-card (adequate for inference) + 64 lanes already covers 6 GPUs with lanes to spare for NVMe/chipset. + +**Lane budget for 6 GPUs on sTRX4 (64 lanes), estimated (no vendor spec gives a topology this +specific — treat this bullet as an estimate):** typical sTRX4 boards reserve ~4 lanes for the +chipset uplink and commonly wire 1-2 M.2 slots directly to the CPU (4 lanes each) — so realistic +GPU-available lanes land around 44-52 of the 64, i.e. **6 GPUs at x8 electrical each (48 lanes) is +plausible but board-model-dependent**; x16-each for 6 cards is not possible on 64 lanes regardless of +board. x8 electrical is not a meaningful inference-speed penalty for llama.cpp (weights are loaded +once; the ongoing per-token traffic across PCIe is small compared to compute), so this is an +acceptable tradeoff, not a real bottleneck for this workload. + +--- + +## 6. Power budget vs. the ~200W / one spare 8-pin headroom + +Official/vendor TDPs: + +| Card | TDP | Source | +|---|---|---| +| GTX 1080 (owned) | 180W, one 8-pin | [confirmed 180W, PCIe 3.0 x16, 1× 8-pin](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu) — spec matches NVIDIA's own launch figures reported across multiple outlets incl. Tom's Hardware | +| RTX 2080 (owned) | 215W | Cross-checked across gpuzoo/cputronic/notebookcheck spec pages, consistent at 215W | +| GT 710 (owned) | ~19W, **no external power connector** (slot power only) | [MSI/EVGA/Zotac GT 710 spec pages](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification) | +| RTX 3060 12GB (candidate purchase) | 170W, one 8-pin | [NVIDIA-confirmed 170W TDP, one 8-pin connector](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/) | +| RTX 3090 24GB (candidate purchase) | 350W, two 8-pin, [NVIDIA's own RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/) lists 350W and a 750W PSU minimum | NVIDIA official | +| R9700 32GB (already deployed / "more of the same") | 300W (per this repo's `server-planing.md`, consistent with AMD's own R9700 product page framing it as a 300W-class card) | in-repo prior research | + +**Against the stated ~200W / one spare 8-pin budget:** + +- Adding **one RTX 3060 12GB** (170W, one 8-pin) is the *only* candidate in this list that fits the + stated headroom as-is — it uses the one spare connector and stays under 200W. +- Adding the already-owned **GTX 1080** (180W) as the classifier's dedicated card also just barely + fits (180W ≤ 200W, one 8-pin) — this is a genuinely free option since the card is already owned and + its power draw is within budget, unlike every purchase candidate below. +- Adding the already-owned **RTX 2080** (215W) **exceeds** the stated 200W headroom by 15W — technically + over budget on paper, though real-world draw is usually a bit under rated TDP; flag it as marginal, + not safely fitting. +- Adding a **second R9700** (300W) or an **RTX 3090** (350W, needs two 8-pin — the user has only one + spare) both blow well past the current power budget on both watts and connector count. +- **The GT 710 draws no meaningful power (~19W, no PCIe power connector at all)** — it is free from a + power-budget standpoint regardless of what else is added. + +**PSU upgrade trigger:** the very first stage that adds *any* GPU beyond a GTX 1080-class card (180W, +one 8-pin) or an RTX 3060 12GB (170W, one 8-pin) exhausts the stated headroom. Any stage that reaches +for a second 32GB-class card (R9700 or equivalent) or any 300W+ card **requires a PSU upgrade before +that stage**, not after — see the roadmap table in §7 for exactly which stage that is. + +--- + +## 7. Mixed-GPU feasibility: ROCm + CUDA, and is the GT 710 usable at all + +**ROCm and CUDA are different llama.cpp builds, but that's exactly the pattern already in this +repo.** `ghcr.io/ggml-org/llama.cpp` publishes both `server-rocm` and `server-cuda` as separate, +independently-built image tags (confirmed present on the [ggml-org container registry](https://github.com/orgs/ggml-org/packages/container/llama.cpp) +and documented in [`docs/docker.md`](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md) +— *"server-cuda: Same as `server` but compiled with CUDA support"*, *"server-rocm: Same as `server` +but compiled with ROCm support"*). You cannot mix backends inside one process/container, but you +**can** run one `server-rocm` container pinned to the R9700 and a separate `server-cuda` container +pinned to an NVIDIA card, simultaneously, on the same host — this is architecturally identical to +today's `llama-server` + `qwen-classifier` two-container split in `docker-compose.yml`, just with a +different image tag for the NVIDIA-backed service and NVIDIA's container runtime (`nvidia-container-toolkit` ++ `--gpus` / device reservation, the CUDA-world equivalent of this repo's `/dev/kfd`+`/dev/dri`+ +numeric-GID ROCm pattern documented in +[`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md)). None of that doc's ROCm-specific +findings (the `GPU_MAX_HW_QUEUES=1` MES firmware workaround, the numeric-GID `group_add` fix) apply to +an NVIDIA/CUDA container — those are ROCm-stack-specific bugs, not general multi-GPU-container issues. + +**Is this an implicit AMD→NVIDIA rebuild, or additive?** Worth surfacing explicitly since the two +source plans conflict on this: `server-planing.md` is an AMD-only, ROCm-only 4×R9700 plan. This +ticket's plan is **additive/mixed** — keep the R9700 running the main model under ROCm, and bolt on +NVIDIA cards under CUDA for secondary roles (classifier, or a second inference GPU for the big model +if going the "more of the same type" route means buying NVIDIA instead of more R9700s). Both are +internally consistent, but they are different end-states — flag this choice back to the user rather +than assuming one. + +**Splitting the *main* 27B model itself across mixed AMD+NVIDIA silicon in one process is not +possible** — llama.cpp's multi-GPU tensor-split only works within a single backend build. To use +both an R9700 and an NVIDIA card for the *same* model's layers, all the compute-hosting cards need to +be the same backend (all-ROCm or all-CUDA) in that one process. This is why §5's roadmap treats "add +GPU capacity to the main model" and "add a GPU for the classifier" as separable purchases with +different backend constraints, not a single mixed pool. + +**Is the GT 710 usable for anything in this pipeline? No.** Reasoning: + +- 1GB VRAM cannot hold any meaningful fraction of either model's weights (17.6 GB / 2.4-4.3 GB) — + even a handful of transformer layers at Q4 quantization exceeds 1GB. +- It's Kepler-generation silicon (192 CUDA cores, no tensor cores) — llama.cpp's CUDA backend + technically supports pre-Turing cards, but at this VRAM size there's nothing to usefully offload. +- It draws power from the PCIe slot only, no external connector — genuinely free to keep installed. +- **Plausible actual use: dedicate it as the box's display-output card**, so every compute-capable + GPU (R9700, and whichever NVIDIA cards get added) can be fully headless/compute-only with none of + their VRAM or a display output tied up driving a monitor — a real, if minor, use for it. This is + this doc's own inference from the spec facts above, not a claim found in any primary source. + +--- + +## 8. GPU market pricing vs. the $20-30/GB assumption + +| Card | VRAM | Backend | Current used-market price (estimate — see caveat) | $/GB | +|---|---|---|---|---| +| RTX 3060 12GB | 12GB | CUDA | ~$240-300 used (eBay listings, [gpupoet.com tracker](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060): *"from $239"*, [eBay live listings](https://www.ebay.com/shop/rtx-3060-12gb) averaging ~$488 asking but with a $239 floor) | **~$20-25/GB** — matches the stated assumption | +| RTX 3090 24GB | 24GB | CUDA | ~$1,010-1,050 used ([bestvaluegpu.com Sep 2026 tracker](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/), [xda-developers coverage](https://www.xda-developers.com/used-rtx-3090-still-best-for-local-ai-in-value/)) | **~$42-44/GB** — well above the stated assumption | +| R9700 32GB ("more of the same type") | 32GB | ROCm | **New only — $1,299 MSRP**, street price $1,400-1,585 as of this research ([overclock3d](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/), [pricehistory.app tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm)) — too recent a release (2026) for a used market to exist yet | **~$41-50/GB, and not a used-market price at all** | + +**Caveat on all three price figures:** these are live marketplace asking-price snapshots pulled via +web search on 2026-09-11, not sold-price data or a vendor spec sheet — treat as directional, not +exact. eBay asking prices in particular run above realized sale prices. + +**Correction to the user's stated assumption:** $20-30/GB is a good estimate specifically for +**last-generation mainstream used cards** (RTX 3060 12GB fits it almost exactly) but **not** for +high-VRAM flagship cards like the RTX 3090 (~1.5-2× that rate) or for "more of the same type" R9700 +units, which aren't used-market at all yet and sit even higher per GB than the 3090. If the plan is +"cheapest path to more VRAM," multiple RTX 3060 12GB cards (or similar mid-tier used cards) beat one +RTX 3090 on $/GB, at the cost of needing more PCIe slots and more total wattage/connectors to reach +the same aggregate VRAM — which is exactly the tradeoff the Threadripper lane-count plan in §5 is +for. + +--- + +## 9. Step-by-step roadmap + +All "resulting max context" figures assume `--parallel 2` and the KV precision stated; per §2, the +`--ctx-size` value shown is the *total* (pre-division) value to pass to llama-server. + +| Stage | Hardware change | Est. cost | Backend | Usable VRAM (main-model pool) | Max context @ parallel=2 (`q4_0` KV) | PSU upgrade triggered? | +|---|---|---|---|---|---|---| +| **0 (current)** | 1× R9700 32GB, in production | $0 | ROCm | 32GB (shared with classifier) | ~131K/slot today at `q8_0` KV (262144 total ÷ 2) | No | +| **1 — classifier isolation** | Move classifier onto the already-owned **GTX 1080** (180W, own container, `server-cuda`), freeing the R9700 entirely for the main model. Matches the existing dual-model pattern qwen-code's own docs describe (§10) and this repo's `qwen-classifier` service already implements, just on separate silicon instead of a shared card. | $0 (already owned) | ROCm (main) + CUDA (classifier) | R9700's full 32GB now available to the main model alone | ~262K/slot @ `q8_0` (unchanged ctx-size, no more classifier contention) | **No** — 180W GTX 1080 fits the stated ~200W/one-8-pin headroom | +| **2 — second big-model GPU** | Add **one more 32GB-class card** for the main model. Cheapest correct-backend option: a second R9700 (~$1,300-1,585 new, ROCm, same backend as the first — required if tensor-splitting one model across two cards) | ~$1,300-1,585 | ROCm | 64GB pooled | `--ctx-size 500000 --parallel 1` fits at `q4_0` (~33GB) or `q8_0` (~48GB, tight but fits in 64GB) — **not yet 500k×2** | **Yes** — 300W card, no spare 8-pin left after stage 1 | +| **3 — reach 500k × 2 / 1M stretch** | No further hardware if stage 2's 64GB pool is used with `--cache-type-k/v q4_0`: `--ctx-size 1000000 --parallel 2` needs ~33-39GB (§4), fits inside 64GB with real headroom for the compute buffer. If `q8_0` KV is required instead (this repo's current quality bar for the main model), the ~51-54GB need is tight-to-marginal on 64GB — a **third** 32GB card (~96GB pool) removes the risk. | $0 (reuses stage 2) or +$1,300-1,585 for a 3rd card if `q8_0` KV is required | ROCm | 64GB (q4_0 case) or 96GB (q8_0 case) | **500k×2 parallel achieved**, and the 1M stretch goal is the *same config* with `--parallel 1` instead of 2 (§2) | Already upgraded at stage 2 | +| **4 — optional CPU/lane platform swap** | Only needed if the plan is to keep scaling past 2-3 big cards, or to add several small used cards (RTX 3060 12GB) for extra headroom/throughput rather than raw ctx-size. Swap to **non-PRO Threadripper 3000-series (sTRX4)** — 64 PCIe 4.0 lanes, ~x8-per-slot for up to 6 GPUs (§5). Threadripper PRO 3000WX/5000WX (128 lanes) only if x16-per-card matters or 6+ full-bandwidth slots are wanted. | Used sTRX4 CPU+board: roughly $400-800 combined on the used market (not independently priced in this pass — **estimate**, not cited) | n/a (platform only) | n/a | n/a | Independent of GPU wattage — driven by whatever GPU count/wattage stage 5+ adds | +| **5+ — scale-out via small used cards** | Add RTX 3060 12GB units (~$20-25/GB, the assumption that actually holds, §8) instead of more 32GB flagship cards, once lane count (stage 4) supports it — useful for extra parallel slots / throughput beyond the 500k×2 target rather than for raising ctx-size further (500k×2/1M is already met at stage 3). | ~$240-300/card | CUDA (separate container per §6) | +12GB pooled per card, but on a *different backend* from the ROCm main model — usable for extra classifier/small-model capacity or a separate CUDA-backend llama-server instance, not as additional tensor-split VRAM for the ROCm main model | Unchanged for the main model; adds parallel capacity elsewhere | Yes, cumulative — each additional 170W card needs PSU headroom stage 2 already consumed | + +**Where the existing dual-model pattern sits in this roadmap:** it's stage 1, and it's free. The +qwen-code docs pattern (main model + a small, always-resident, non-thinking fast/classifier model — +see §10) is already implemented in this repo; the only roadmap-relevant change is *which GPU* the +classifier sits on, moving it off the R9700 entirely onto an already-owned NVIDIA card frees the +R9700's full 32GB for the 500k×2/1M push instead of splitting it with the classifier as happens +today. + +### 9.1 Upgrade path, as diagrams + +Diagram form of the same §9 table and §11.9's dense-vs-Flash-Next call — nothing new is claimed here, +this is a visual index back into the cited sections above. + +**Stage-by-stage hardware path** (PSU-upgrade triggers and target reached called out inline): + +```mermaid +flowchart TD + S0["Stage 0 — today
1x R9700 32GB, ROCm
classifier shares the card
$0"] + S1["Stage 1 — classifier isolation
+ GTX 1080 (owned, 180W, CUDA)
R9700 freed for main model
$0 · PSU OK (180W fits ~200W headroom)"] + S2["Stage 2 — 2nd big-model GPU
+1x R9700 32GB (ROCm)
64GB pooled
~$1,300-1,585 · PSU UPGRADE REQUIRED (300W, no 8-pin left)"] + S3q4["Stage 3a — q4_0 KV
--ctx-size 1,000,000 --parallel 2
~33-39GB, fits in 64GB
$0 (reuses stage 2)"] + S3q8["Stage 3b — q8_0 KV (current prod quality)
~51-54GB, tight on 64GB
+1x R9700 -> 96GB removes risk
+~$1,300-1,585"] + TARGET(["500k x2 parallel reached
= 1M stretch goal, same VRAM
(--parallel 1 vs 2 is a config flag, §2)"]) + S4["Stage 4 — platform swap (optional)
sTRX4 Threadripper 3000, 64 PCIe4 lanes
only needed past 2-3 big cards
~$400-800 (estimate, §9)"] + S5["Stage 5+ — scale out
+RTX 3060 12GB cards (CUDA, separate backend)
extra parallel/throughput, not more ctx-size
~$240-300/card · PSU upgrade each card"] + + S0 --> S1 --> S2 + S2 --> S3q4 --> TARGET + S2 --> S3q8 --> TARGET + TARGET -.->|"only if scaling past this"| S4 --> S5 + + style TARGET fill:#2e7d32,color:#fff,stroke:#1b5e20 + style S2 fill:#8a5a00,color:#fff,stroke:#5c3d00 + style S5 fill:#8a5a00,color:#fff,stroke:#5c3d00 +``` + +**Model choice, and the one open question that could change it** (§11.9): + +```mermaid +flowchart TD + Q{"Goal: 500k-1M ctx
within a $20-30/GB VRAM budget?"} + D["Dense Qwen3.8-27B
17.6-54.7GB weights (Q4_K_XL-BF16)
reaches 500k x2 on 2-3 cards,
500k x4 on 2-6 cards depending on quant
(§11.7 tables)"] + F{"Try --n-cpu-moe:
offload MoE experts to system RAM?
(untested for this model, §11.8)"} + FBAD["Flash-Next, all-GPU weights
111-354GB just for weights
needs 4-13 cards before any KV cost
NOT recommended at this budget (§11.9)"] + FGOOD["Flash-Next, experts in system RAM
GPU VRAM could shrink a lot
2.67x cheaper KV/token becomes relevant
UNVERIFIED — prototype on real server first"] + CAVEAT["+ real caveat either way:
PR #27742 flags unverified conv branch,
3% QSA divergence, prefill-pos-0-only PLE
(§11.1) — dense model carries no such flag"] + + Q --> D + Q -->|"considering Flash-Next instead"| F + F -->|"works well"| FGOOD + F -->|"doesn't help / untested"| FBAD + FGOOD --> CAVEAT + FBAD --> CAVEAT + + style D fill:#2e7d32,color:#fff,stroke:#1b5e20 + style FBAD fill:#8a1c1c,color:#fff,stroke:#5c1212 + style FGOOD fill:#8a5a00,color:#fff,stroke:#5c3d00 +``` + +--- + +## 10. Qwen-code's own docs on the fast-model/classifier pattern + +Fetched directly per the user's link: +[qwenlm.github.io/qwen-code-docs/en/users/overview/](https://qwenlm.github.io/qwen-code-docs/en/users/overview/) +— **the overview page itself does not describe the dual-model/classifier pattern**; it only covers +single-model-provider setup (Alibaba ModelStudio / third-party / custom provider), one model at a +time. The actual fast-model/classifier documentation lives on the **Auto Mode** page instead, which +this repo's own `fast-model-choice.md` already fetched and cited in detail: +[qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/) — +summary (see `fast-model-choice.md` §1 for the full quote): a two-stage classifier gate, both stages +using "your configured fast model (`/model --fast`)", Stage 1 a ~300ms `{shouldBlock}`-only check, +Stage 2 a ~3-5s chain-of-thought review that only runs on a Stage-1 block. Nothing in either page +gives a recommended *context size* or *model size* for the fast model beyond what's implied by that +latency budget — this repo's own prior research (`fast-model-choice.md`) derived the actual context +requirement from qwen-code's source code instead (`packages/core/src/permissions/classifier-transcript.ts`), +since the docs pages don't state one. No new information changes that prior doc's conclusion; this +section exists to confirm the overview page was checked directly as instructed and doesn't contradict +or add to it. + +--- + +## 11. Alternative: Qwen3.8-Flash-Next (MoE, hybrid attention) + +The user also wants to weigh switching (or adding) **Qwen3.8-Flash-Next** — a 125B-total/6B-active MoE +with a hybrid recurrent-attention architecture — against staying on dense Qwen3.8-27B, and separately +wants this section to cover **going up in weight quant** (Q4_K_XL → Q6_K_XL → Q8_0 → BF16/fp16) for +*both* models, not just Q4. Feasibility first, since it gates everything else. + +### 11.1 Feasibility verdict: supported, but immature — read before trusting any number below + +Checked directly against the primary sources the task named: + +- **llama.cpp mainline support exists.** [PR #27742](https://github.com/ggml-org/llama.cpp/pull/27742) + ("model: add Qwen3.8-Flash-Next (qwen4exp)") was **merged into `master` on 2026-08-27** by ngxson. + It adds the full architecture: Gated DeltaNet layers (sigmoid-gated linear attention), QSA + ("Qwen Sparse Attention", operating at micro-block granularity), hyper-connections, and the PLE + n-gram embedding table. `llama.cpp`'s own docs list CPU/CUDA/Metal/ROCm as supported backends for + it — this is not a CUDA-only feature. +- **This repo's pinned image is a floating tag, not a version pin.** `docker-compose.yml` runs + `ghcr.io/ggml-org/llama.cpp:server-rocm` with no date/digest suffix — a rolling "latest ROCm server + build" tag, not a release version. The merge is from 2026-08-27, and today is 2026-09-11 (~2 weeks + later), so a **fresh pull** of `server-rocm` should include it — but whatever image is already + cached/running on the R9700 box may predate the merge. **Action before touching this model on the + server: `docker compose pull llama-server` and check the startup log's build/commit banner is dated + on/after 2026-08-27**, not just "the tag says server-rocm." +- **Real, primary-source-flagged immaturity — this is the part that should temper enthusiasm.** The + PR's own description/review discussion states: *"The conv branch itself is still numerically + unverified because the fixture zeroes its weights"*; QSA sparse attention *"diverges on 3 percent of + positions"* above its budget threshold; the PLE depthwise convolution *"is exact only for a prefill + that starts at position 0"* (i.e. correctness is not guaranteed once `--cache-reuse`/prompt-caching + is in play — a flag this repo already turns on for the dense model per the latest commit). None of + that is disqualifying, but it is a primary-source admission that this is a fresh, not-fully-verified + implementation, not a mature, widely-battle-tested one like the dense Qwen3.8-27B path. +- **Multi-slot serving needs an explicit new flag.** The same PR states: *"`set_input_qsa` asserted + `n_stream == 1`, so llama-server could not serve this model with more than one slot unless `-kvu` + was passed."* Per the server README (§2), `--kv-unified`/`-kvu` defaults to enabled **only when slot + count is auto** (`-1`). This repo's compose file sets `--parallel ${LLAMA_PARALLEL:-2}` **explicitly** + (not auto) — so adopting Flash-Next with `--parallel` > 1 requires **adding `--kv-unified` (or + `-kvu`) to the launch flags**, a real deploy-time change, not something that "just works" by copying + today's flag set onto a new model file. + +**Verdict: yes, runnable** on this repo's backend (ROCm, mainline, no dev branch needed) as long as the +image is pulled after 2026-08-27 and `-kvu` is added for multi-slot use — but treat it as +**usable-with-caution**, not a drop-in swap, given the PR author's own unresolved-correctness notes. + +### 11.2 Architecture, verified against `config.json` directly + +Fetched from `Qwen/Qwen3.8-Flash-Next`'s `config.json` (unsloth's GGUF repo repackages the same base +model): `num_hidden_layers=48`, `hidden_size=2560`, `num_attention_heads=24`, `num_key_value_heads=2`, +`head_dim=256`, `max_position_embeddings=262144` (same native/extensible-to-1M framing as the dense +model — same YaRN quality caveat from §3 applies here too, unverified past native length), `num_experts=512`, +`num_experts_per_tok=10`, and the linear-attention head config: `linear_num_key_heads=16`, +`linear_num_value_heads=48`, `linear_key_head_dim=128`, `linear_value_head_dim=128`. + +Layer pattern (confirmed both from the model card's own description and `config.json`'s +`full_attention_interval=4`): every 4th layer is full/QSA attention, the other 3 are Gated DeltaNet — +**12 of 48 layers grow a real KV cache; the other 36 have a fixed-size recurrent state that does not +grow with context length.** (24% full-attention layers vs. the dense model's 16-of-64 = 25% — similar +ratio, but the *absolute* per-layer KV cost differs because `num_key_value_heads` is 2 here vs. 4 on +the dense model — see below.) + +### 11.3 Per-token growing-KV-cache cost + +``` +12 full-attention layers × 2 (K+V) × 2 kv_heads × 256 head_dim × 2 bytes (fp16) = 24 KiB/token +``` + +| Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 | +|---|---|---|---| +| 262,144 (native) | ~6.0 GiB | ~3.0 GiB | ~1.5 GiB | +| 500,000 | ~11.4 GiB | ~5.7 GiB | ~2.9 GiB | +| **1,000,000 (500k×2, or 1M stretch)** | **~22.9 GiB** | **~11.4 GiB** | **~5.7 GiB** | +| **2,000,000 (500k×4)** | **~45.8 GiB** | **~22.9 GiB** | **~11.4 GiB** | + +`--cache-type-k/v` are the same generic llama.cpp KV-cache-quantization flags used elsewhere in this +doc; nothing in the PR or the server README suggests they're handled differently for the 12 +full-attention layers of a hybrid model — they quantize the same growing K/V buffers as on a plain +transformer. (No primary source explicitly confirms this for *this* architecture specifically — flagged +as a reasonable extrapolation, not a directly-cited fact, same caveat class as this doc's other +estimates.) + +### 11.4 Fixed (non-growing) recurrent state — Gated DeltaNet layers + +The 36 Gated DeltaNet layers each keep a fixed-size recurrent state (an outer-product-style +key×value matrix per head) that does **not** scale with context length — only with slot/sequence +count. Sized from `config.json`'s linear-attention head params: + +``` +36 layers × linear_num_value_heads(48) × linear_key_head_dim(128) × linear_value_head_dim(128) × 4 bytes (fp32 state) +≈ 36 × 48 × 128 × 128 × 4 bytes ≈ 108 MiB per slot +``` + +This is **this doc's own derivation from the published head-dimension params, not a value pulled +directly from llama.cpp source or docs** — the PR text confirms the state exists per-stream/per-slot +but doesn't publish an exact byte formula, so treat the ~108 MiB/slot figure as an estimate, medium +confidence. Even at 4 parallel slots that's under half a gigabyte — **negligible** next to both the +growing KV cache (GBs) and the weights (tens to hundreds of GB) computed below. The headline +implication holds regardless of the exact multiplier: Flash-Next's "big memory line item" is the MoE +weights, not the attention state of any kind. + +### 11.5 Magnitude vs. the dense model — how much cheaper is KV, really + +At the same total ctx-size, Flash-Next's growing KV cache is **24 KiB/token vs. the dense model's +64 KiB/token — 2.67× smaller**, i.e. Flash-Next's KV budget is **37.5%** of the dense model's at +identical context length. This is a real, significant win *for the KV-cache line item specifically* — +but see §11.9: it's a much smaller slice of a much bigger total, because the weights move the other +way by a far larger factor. + +### 11.6 Weight sizes — verified from each unsloth GGUF repo's actual file listing + +Fetched directly from the HF file trees (not estimated from ratios), current as of this research pass: + +| Quant tier | Qwen3.8-27B (dense) | Qwen3.8-Flash-Next (MoE) | +|---|---|---| +| Q4_K_XL (`UD-Q4_K_XL`) | **17.6 GB** (existing baseline) | **111.4 GB** (4 parts: 10.9MB + 49.9GB + 49.4GB + 12.1GB) | +| Q6_K_XL (`UD-Q6_K_XL`) | **25.3 GB** | **169 GB** (6 parts) | +| Q8_0 | **29 GB** | **188 GB** (6 parts) | +| BF16/fp16 | **54.67 GB** (50GB + 4.67GB, 2 parts) | **354 GB** (8 parts) | + +Sources: [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main) +and its `BF16/` subfolder; [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main) +and its `UD-Q4_K_XL/`, `UD-Q6_K_XL/`, `Q8_0/`, `BF16/` subfolders (per-file sizes summed). The +preliminary Q8_0 figure floated before this research pass (~192GB) was slightly high — the real +listing sums to **188 GB**; everything else in the preliminary list was accurate to within rounding. + +**The weight-quant axis and the KV-cache-quant axis are independent knobs.** Raising weight quality +(Q4_K_XL → BF16) does not require raising `--cache-type-k/v` — the two flags are unrelated, and this +repo already proves that pattern works (`q8_0` KV cache is deployed today against `Q4_K_XL` weights). +A user chasing **maximum output quality** can run e.g. **BF16 weights + `q4_0` KV cache** — full-precision +weights for quality, still-compressed KV for context budget — or any other combination in the tables +below; nothing about picking a higher weight quant forces a matching KV precision. + +### 11.7 Total VRAM: does it fit, across quant tiers and both parallelism targets + +All totals = weights + growing KV cache + an estimated **+3-6 GB** compute-buffer/runtime overhead +(same estimate band as §4, carried over — not re-derived for this architecture; flagged medium +confidence there too). "Cards" = ceil(total ÷ 32GB), i.e. how many R9700-class 32GB cards it takes. + +#### Dense Qwen3.8-27B — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`) + +| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) | +|---|---|---|---| +| Q4_K_XL (17.6GB) | ~79-85GB (**3**) | ~48-54GB (**2**) | ~33-39GB (**2**) | +| Q6_K_XL (25.3GB) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) | ~44-47GB (**2**) | +| Q8_0 (29GB) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) | ~47-50GB (**2**) | +| BF16 (54.67GB) | ~119-122GB (**4**) | ~88-91GB (**3**) | ~73-76GB (**3**) | + +#### Dense Qwen3.8-27B — 500k × 4 parallel (`--ctx-size 2,000,000`) + +| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) | +|---|---|---|---| +| Q4_K_XL (17.6GB) | ~143-146GB (**5**) | ~82-85GB (**3**) | ~51-54GB (**2**) | +| Q6_K_XL (25.3GB) | ~150-153GB (**5**) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) | +| Q8_0 (29GB) | ~154-157GB (**5**) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) | +| BF16 (54.67GB) | ~180-183GB (**6**) | ~119-122GB (**4**) | ~88-91GB (**3**) | + +#### Qwen3.8-Flash-Next — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`) + +| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) | +|---|---|---|---| +| UD-Q4_K_XL (111.4GB) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) | ~120-123GB (**4**) | +| UD-Q6_K_XL (169GB) | ~195-198GB (**7**) | ~183-186GB (**6**) | ~178-181GB (**6**) | +| Q8_0 (188GB) | ~214-217GB (**7**) | ~202-205GB (**7**) | ~197-200GB (**7**) | +| BF16 (354GB) | ~357-360GB (**12**) | ~357-360GB (**12**) | ~357-360GB (**12**) | + +#### Qwen3.8-Flash-Next — 500k × 4 parallel (`--ctx-size 2,000,000`) + +| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) | +|---|---|---|---| +| UD-Q4_K_XL (111.4GB) | ~160-163GB (**6**) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) | +| UD-Q6_K_XL (169GB) | ~218-221GB (**7**) | ~195-198GB (**7**) | ~183-186GB (**6**) | +| Q8_0 (188GB) | ~234-237GB (**8**) | ~211-214GB (**7**) | ~199-202GB (**7**) | +| BF16 (354GB) | ~397-400GB (**13**) | ~377-380GB (**12**) | ~366-369GB (**12**) | + +(Flash-Next's KV precision barely moves the total at any weight quant above `UD-Q6_K_XL` — the weights +so dominate the budget that KV quantization stops mattering for the "how many cards" question. This +is the clearest signal in this whole section: for Flash-Next, the weight-quant choice is the entire +hardware-sizing decision; for the dense model, KV precision still matters a lot.) + +### 11.8 CPU MoE-expert offload — the one lever that could change this calculus + +Flash-Next is a 512-expert/10-active-per-token MoE, and llama.cpp has a purpose-built flag for exactly +this shape of model, confirmed directly from the server README: **`--n-cpu-moe`** — *"keep the Mixture +of Experts (MoE) weights of the first N layers in the CPU"* — plus the more general +**`--override-tensor`** (*"override tensor buffer type"*, pattern-matched by tensor name) that the same +flag is built on top of. Both are generic, architecture-agnostic llama.cpp mechanisms (they match on +tensor name patterns, not model type), so there's no reason to expect them not to apply to Flash-Next's +MoE tensors specifically — but this pass found **no primary source that has actually tested +`--n-cpu-moe` against this specific qwen4exp architecture**, so treat "it works here" as plausible, +not confirmed. + +If it does work as expected, this changes the whole weight-VRAM picture in §11.7: the ~90-95% of +Flash-Next's weight footprint that's MoE expert tensors could live in system RAM while attention +projections, the shared/non-expert tensors, and the full KV cache stay on GPU — meaning a much smaller +GPU-VRAM number than the "all weights on GPU" tables above, at the cost of PCIe/RAM-bandwidth-bound +inference speed for whichever experts get selected per token (this repo has no benchmark of that +tradeoff, and it's highly system-RAM-bandwidth-dependent, so no number is given here — flagged as an +escape hatch worth prototyping directly on the server, not something this research values responsibly +without a real test run). + +### 11.9 Net recommendation: dense Qwen3.8-27B vs. Flash-Next, for this user's stated goal + +**Net loss for this user's goal, as things stand — stay on dense Qwen3.8-27B.** Reasoning: + +- The user's target (500k×2 or 500k×4, on a $20-30/GB-VRAM budget, GPUs in 32GB increments) is a + **VRAM-budget-constrained** goal, and §11.7 shows Flash-Next's *weights alone* (111-354GB depending + on quant) dwarf the entire dense-model total-VRAM figure from §4/§11.7 (33-183GB depending on quant) + at every parallelism target. Flash-Next's much cheaper per-token KV cache (§11.5, real and verified) + is a rounding error next to that weight-size gap — the "2.67× cheaper KV" win doesn't come close to + offsetting a "6-20× larger weight footprint," so at $20-30/GB-VRAM the *dense* model reaches 500k×2 + or 500k×4 for a fraction of the card count and dollar cost that Flash-Next needs even at its lowest + usable quant (`UD-Q4_K_XL`, 4-5 cards minimum) — before even factoring in §11.1's immaturity flags. +- The one scenario that could flip this verdict is `--n-cpu-moe` actually working well for this + architecture (§11.8) — if most of those 111-354GB of expert weights can sit in system RAM at + acceptable throughput, Flash-Next's GPU-VRAM number could shrink dramatically and its real KV-cache + advantage would start to matter. That is untested here and shouldn't be assumed; it's the one + concrete next step worth trying on the actual server before ruling Flash-Next out permanently. +- Independent of VRAM: §11.1's primary-source-flagged correctness caveats (unverified conv branch, + 3%-divergence QSA, prefill-position-0-only PLE exactness) are a real quality/stability risk on a + production coding-agent stack that dense Qwen3.8-27B simply doesn't carry, since it's been running + in this repo already. + +--- + +## 12. 4-parallel × 500k scenario — all four combinations side by side + +Per §2's already-established, cited rule (`n_ctx_per_seq = n_ctx / n_parallel`, +[ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681)), the same division +applies at 4 slots: **500k tokens on each of 4 parallel slots requires `--ctx-size 2,000,000`** — +double the 2-parallel target's `--ctx-size 1,000,000`, for the same reason 500k×2 needed double +262,144. This isn't a new mechanism, just the same formula at `--parallel 4`. + +Full per-quant-tier tables for all four combinations are in §11.7 above (dense×2, dense×4, Flash-Next×2, +Flash-Next×4 are each their own table there). Headline comparison at the KV precision already proven +in production in this repo (`q8_0`) and each model's respective current/cheapest-usable weight quant: + +| Scenario | `--ctx-size` | Weight quant | q8_0-KV total VRAM | Cards (32GB) | +|---|---|---|---|---| +| Dense × 2 (or 1M stretch) | 1,000,000 | Q4_K_XL (17.6GB, current) | ~48-54GB | **2** | +| Dense × 4 | 2,000,000 | Q4_K_XL (17.6GB, current) | ~82-85GB | **3** | +| Flash-Next × 2 (or 1M stretch) | 1,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~126-129GB | **4**, edge | +| Flash-Next × 4 | 2,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~137-140GB | **5** | + +**4-parallel × 500k reachability against this repo's existing roadmap stages (§9):** + +- **(a) Current 1×R9700 32GB:** none of the four combinations fit — not even dense×2 at any weight/KV + quant (§4's own conclusion, unchanged). +- **(b) The 2-3×R9700 roadmap already proposed in §9 (64-96GB):** covers **dense×2 fully** (stage 3, as + already established) and **dense×4 at `q4_0` KV with Q4_K_XL or Q6_K_XL weights** (~51-62GB, fits in + 64-96GB) — but **not** dense×4 at higher weight quants (Q8_0/BF16 need 3-6 cards depending on KV + precision, per §11.7's dense×4 table) and **not any Flash-Next scenario** (minimum is 4 cards/128GB + even at the cheapest usable quant and tightest KV). +- **(c) The full 4-6×R9700 stretch scenario** (`server-planing.md`'s original plan, 128-192GB pooled): + covers **dense×4 at every weight quant up to BF16** (worst case ~91GB at BF16+q4_0, well inside + 128GB) and **Flash-Next×2 at `UD-Q4_K_XL`** (126-140GB, fits a 5-card/160GB build, tight on a 4-card/ + 128GB one) — but **not** Flash-Next×4 at any weight quant above `UD-Q4_K_XL`, and not Flash-Next at + `BF16` under any parallelism (needs 12-13 cards, an entirely different scale of build than anything + in this doc's roadmap). + +--- + +## Sources + +- [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) — `--ctx-size`, `--parallel`, `--cache-type-k/v`, `--kv-unified`, `--cache-reuse`, `--n-cpu-moe`, `--override-tensor` flag definitions +- [ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681) — real server log proving `n_ctx_per_seq = n_ctx / n_parallel` +- [ggml-org/llama.cpp#27742](https://github.com/ggml-org/llama.cpp/pull/27742) — "model: add Qwen3.8-Flash-Next (qwen4exp)", merged 2026-08-27; architecture details, `n_stream == 1` / `-kvu` multi-slot requirement, and the conv-branch/QSA-divergence/PLE-prefill correctness caveats +- [Qwen/Qwen3.8-27B config.json](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json) +- [Qwen/Qwen3.8-Flash-Next config.json](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/raw/main/config.json) +- [Qwen/Qwen3-4B-Instruct-2507 config.json](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json) +- [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main) — 17.6GB (Q4_K_XL), 25.3GB (Q6_K_XL), 29GB (Q8_0), 54.67GB (BF16) weight sizes +- [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main) — 111.4GB (UD-Q4_K_XL), 169GB (UD-Q6_K_XL), 188GB (Q8_0), 354GB (BF16) weight sizes, summed from each quant's per-file listing +- [ggml-org/llama.cpp docs/docker.md](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md) — `server-cuda`/`server-rocm` separate image tags +- [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html) +- [Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2) +- [PCWorld — Threadripper PRO launch, 128 PCIe lanes](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html) +- [NVIDIA — GeForce RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/) +- [Lowyat.net — RTX 3060 official 170W TDP](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/) +- [BuildMyServer — GTX 1080 180W/PCIe3.0/1×8-pin spec listing](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu) +- [MSI — GT 710 1GD5 LP spec page](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification) +- [overclock3d — AMD Radeon AI PRO R9700 $1,299 MSRP](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/) +- [pricehistory.app — R9700 street price tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm) +- [bestvaluegpu.com — RTX 3090 used price tracker, Sep 2026](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/) +- [gpupoet.com — RTX 3060 12GB used listings](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060) +- [Qwen Code docs — overview](https://qwenlm.github.io/qwen-code-docs/en/users/overview/) +- [Qwen Code docs — Auto Mode (fast-model pattern)](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/) +- This repo: [`docker-compose.yml`](../../docker-compose.yml), [`.env.example`](../../.env.example), [`docs/server-planing.md`](../server-planing.md), [`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), [`qwen3.8-27b-tool-calling.md`](qwen3.8-27b-tool-calling.md), [`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md), [`fast-model-choice.md`](fast-model-choice.md) + +## Confidence/uncertainty summary + +- **High confidence:** the KV-cache-per-token formulas for both dense models and Flash-Next (computed + directly from each model's own `config.json`, same method this repo's prior research already used + and cross-checked); the `n_ctx_per_seq = n_ctx / n_parallel` division behavior (directly evidenced by + a real server log in a llama.cpp GitHub issue, and independently already documented in this repo's + own `.env.example`) — and confirmed to apply identically at `--parallel 4` since the mechanism is + parallel-count-agnostic; official TDP figures for GTX 1080, RTX 2080, RTX 3060, RTX 3090, GT 710 + (each cross-checked against 2+ independent spec listings or the vendor's own product page); the + TR4-is-PCIe3/sTRX4-is-PCIe4 generational split (direct launch-coverage quote); the existence of + separate `server-cuda`/`server-rocm` llama.cpp image tags; Qwen3.8-Flash-Next's `config.json` + architecture params and PR #27742's merge date/status and its own stated correctness caveats and + `-kvu` multi-slot requirement (all directly quoted from the primary source); the weight file sizes + for both models at all four quant tiers (summed directly from each HF repo's real file listing, not + estimated). +- **Medium confidence:** the compute-buffer/runtime-overhead estimate in §4/§11.7 (+3-6 GiB) — + extrapolated from one in-repo incident's before/after numbers, not a llama.cpp-documented formula, + and carried over to Flash-Next without re-derivation for its different architecture; the Gated + DeltaNet fixed recurrent-state size in §11.4 (~108 MiB/slot) — this doc's own derivation from the + published head-dimension config, not a value found in llama.cpp source or docs; whether + `--cache-type-k/v` quantization applies identically to Flash-Next's 12 full-attention layers as it + does to a plain transformer (reasonable extrapolation, not directly confirmed for this architecture); + whether `--n-cpu-moe`/`--override-tensor` actually work against Flash-Next's specific MoE tensor + layout (architecture-agnostic mechanism, but untested against this model by any primary source found); + real-world PCIe lane availability for 6 GPUs on a specific sTRX4 board (§5) — no single board's exact + lane map was fetched, this is a reasonable-but-unverified estimate from typical sTRX4 board behavior. +- **Low confidence / explicitly estimated, not cited fact:** all used-GPU marketplace pricing (§8) — + live asking-price snapshots from a single search pass, not sold-price data; the used sTRX4 + CPU+motherboard combo price in the roadmap's stage 4 (§9) — not researched at all in this pass, + flagged as a placeholder estimate; whether YaRN-scaled 500k/1M context actually holds output + quality for either Qwen3.8-27B or Qwen3.8-Flash-Next — no primary source (Qwen's own docs included) + publishes long-context quality benchmarks past the 262,144 native length for either model, so this is + a known-unknown carried forward from each model card's "YaRN-extensible" claim, not a verified + capability; whether the specific `ghcr.io/ggml-org/llama.cpp:server-rocm` image currently cached on + this repo's server actually postdates PR #27742's 2026-08-27 merge — not checked against the live + server in this pass, flagged as an action item in §11.1 rather than a confirmed fact.