diff --git a/docs/research/500k-context-2-parallel-agents-gpu-roadmap.md b/docs/research/500k-context-2-parallel-agents-gpu-roadmap.md
new file mode 100644
index 0000000..8a2ce39
--- /dev/null
+++ b/docs/research/500k-context-2-parallel-agents-gpu-roadmap.md
@@ -0,0 +1,737 @@
+# Research: hardware roadmap to 500k-token context × 2 parallel agents (1M stretch)
+
+**Date:** 2026-09-11
+
+**Question:** What VRAM does 500k-token context × 2 parallel llama-server slots (and a 1M-token
+stretch goal) actually cost for the Qwen3 family, and what hardware roadmap gets there from the
+current single-R9700 setup — given the user's stated plan to add an older (PCIe 4.0) Threadripper
+for lane count, reuse existing RAM/PSU (~200W headroom / one spare 8-pin), mix in already-owned
+NVIDIA cards (GTX 1080 8GB, RTX 2080 8GB, GT 710 1GB) for the classifier role, and price used GPUs
+at roughly $20-30/GB VRAM?
+
+**Answer, short version:** The two goals ("500k × 2 parallel" and "1M stretch") turn out to need
+**the same total VRAM budget** — because of how llama-server's `--ctx-size` and `--parallel` interact
+(§2), 500k × 2 slots and a single 1M-token slot both require setting `--ctx-size 1000000`. At
+`q4_0`-quantized KV cache that's **~33 GB** (weights + KV) for Qwen3.8-27B, at `q8_0` it's **~48 GB**,
+at fp16 it's **~79 GB** — before compute-buffer overhead. That does not fit on the current single
+32GB R9700 at any KV precision, and comfortably fits on two 32GB-class cards only at `q8_0`/`q4_0`.
+The user's $20-30/GB pricing intuition holds for last-gen used consumer cards (RTX 3060 12GB) but
+**not** for RTX 3090 24GB (~$44/GB currently) or a second R9700 (~$41/GB, new — no used market yet
+for a card released mid-2026). The stated Threadripper plan needs to specifically target the
+**non-PRO Threadripper 3000 series on sTRX4** (64 lanes, PCIe 4.0) — older Threadripper on the
+original TR4 socket (1000/2000 series) is PCIe 3.0 only, which doesn't match the user's own PCIe 4.0
+requirement. The power budget (~200W / one spare 8-pin) is exhausted by a *single* mid-tier used GPU
+addition — a PSU upgrade is not optional past the very first stage. See §7 for the roadmap.
+
+---
+
+## 1. Current state (from this repo)
+
+From `docker-compose.yml` and `.env.example` at the repo root:
+
+- **Main model:** `Qwen3.8-27B-UD-Q4_K_XL.gguf` (17.6 GB weights), `--ctx-size 262144`,
+ `--parallel 2`, `--flash-attn on`, `--cache-type-k q8_0 --cache-type-v q8_0`, `--n-gpu-layers 999`,
+ on one AMD Radeon AI PRO R9700 (32GB, ROCm/HIP, `gfx1201`).
+- **Classifier ("fast") model:** `Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`, `--ctx-size 65536`,
+ `--parallel 1`, `--n-gpu-layers 28` (partial offload), `--cache-type-k/v q4_0`, its own container on
+ the *same* R9700, sharing VRAM with the main model — see
+ [`docker-compose.yml`](../../docker-compose.yml) lines ~58-99 and
+ [`fast-model-choice.md`](fast-model-choice.md).
+- `.env.example` already documents the exact fact this research turns on: *"Each slot gets
+ `LLAMA_CTX_SIZE / LLAMA_PARALLEL` tokens of context"* — i.e. today's 262144 ctx-size ÷ 2 parallel
+ slots means each real request only gets **~131K tokens**, not the full 262144, confirmed in-repo
+ before any external source was checked.
+- Prior research already worked out the KV-cache formula for this exact model
+ ([`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md)) — this doc reuses and extends that math for the
+ 500k/1M targets rather than re-deriving it.
+
+`docs/server-planing.md` describes a **different, earlier plan**: a 4× AMD Radeon AI PRO R9700 rig
+on a Gigabyte MZ32-AR0 (single-socket SP3/EPYC, 128 PCIe 4.0 lanes), fully AMD/ROCm. The user's plan
+in this ticket is not that — it pivots toward an older **Threadripper** (SP3's sibling desktop-HEDT
+socket family, not SP3 itself) and explicitly wants to mix in already-owned **NVIDIA** cards. These
+two plans are **not the same build** and, per §6, ROCm and CUDA cards cannot share one llama.cpp
+process — they can only coexist as separate containers on separate cards. Treat `server-planing.md`
+as superseded context, not the active plan, unless the user says otherwise.
+
+---
+
+## 2. llama-server parallelism: does each slot get its own full `--ctx-size`, or is it divided?
+
+**Divided.** This is the single fact that changes the whole budget by 2×, confirmed from three
+independent primary sources:
+
+1. **This repo's own `.env.example`** (quoted above) already documents it for the current deployment.
+2. **llama.cpp's own server README**, fetched directly
+ ([`tools/server/README.md`](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)):
+ `--ctx-size (-c)`: *"size of the prompt context (default: 0, 0 = loaded from model)"*;
+ `--parallel (-np)`: *"number of server slots (default: -1, -1 = auto)"* — the docs list these as
+ independent flags, but don't spell out the division themselves.
+3. **A real user's server log**, quoted verbatim in
+ [ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681), is the actual proof:
+ running with `--ctx-size 327680 --parallel 6` produces `n_ctx = 327680`,
+ `n_ctx_per_seq = 54613` — i.e. `327680 / 6 ≈ 54613`. The reporter explicitly asked for a
+ `--ctx-size-per-seq`-style flag to *avoid* this division; no such flag exists as of the fetch date.
+
+Practical consequence: **to get 500,000 usable tokens on each of 2 parallel slots, `--ctx-size` must
+be set to 1,000,000, not 500,000.** The KV cache is sized off the *total* `--ctx-size`
+(`--kv-unified`, on by default when slots are auto per the README's `-kvu` entry, uses one shared
+pool sized to the full `n_ctx`) — so the VRAM cost of "500k × 2 parallel" and "one 1M-token slot"
+is **identical**: both require `--ctx-size 1000000`. This is a genuinely useful finding for the
+roadmap — reaching the 500k×2 target and the 1M stretch goal cost the same VRAM; the only difference
+is `--parallel 1` vs `--parallel 2` at deploy time, a config change with zero extra hardware cost.
+
+`--cache-type-k` / `--cache-type-v` accept `f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1`
+(default `f16`), per the same README fetch. The repo already uses `q8_0` on the main model and `q4_0`
+on the classifier, so both quantization tiers used in the math below are already-proven-working
+configurations in this stack, not hypothetical flags.
+
+---
+
+## 3. KV-cache math per model
+
+### Qwen3.8-27B (hybrid Gated-DeltaNet / attention)
+
+Reusing the architecture params already pulled from
+[`Qwen/Qwen3.8-27B/config.json`](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json) in
+[`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), re-verified directly for this doc: `num_hidden_layers=64`,
+`full_attention_interval=4` → **16 of 64 layers are standard KV-caching attention**, the other 48 are
+Gated DeltaNet linear-attention layers with a small, context-length-*independent* recurrent state
+(tens of MB total, negligible next to the attention KV cache — ignored below).
+`num_key_value_heads=4` (GQA), `head_dim=256`. Native context `max_position_embeddings=262144`
+(YaRN-extensible to 1M per the model card — **both the 500k and 1M targets exceed native context and
+require RoPE/YaRN scaling**, which is a real quality caveat, not just a memory one — Qwen has not
+published independent long-context quality benchmarks past native length that this research found).
+
+Per-token KV cache, fp16, both K and V, across the 16 full-attention layers:
+
+```
+16 layers × 2 (K+V) × 4 kv_heads × 256 head_dim × 2 bytes = 64 KiB/token
+```
+
+| Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 |
+|---|---|---|---|
+| 262,144 (current) | ~16.0 GiB | ~8.0 GiB | ~4.0 GiB |
+| 500,000 | ~30.5 GiB | ~15.3 GiB | ~7.6 GiB |
+| **1,000,000 (500k×2, or 1M stretch)** | **~61.0 GiB** | **~30.5 GiB** | **~15.3 GiB** |
+
+(`q8_0` is 8-bit vs. fp16's 16-bit → exactly half; `q4_0` is 4-bit → exactly quarter, per llama.cpp's
+own cache-type byte widths.)
+
+### Qwen3-4B-Instruct-2507 (plain GQA transformer, classifier role)
+
+From [`Qwen/Qwen3-4B-Instruct-2507/config.json`](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json)
+(already pulled in [`fast-model-choice.md`](fast-model-choice.md)): `num_hidden_layers=36` — every
+layer is standard attention here (no hybrid split), `num_key_value_heads=8`, `head_dim=128`.
+
+```
+36 layers × 2 (K+V) × 8 kv_heads × 128 head_dim × 2 bytes = 144 KiB/token
+```
+
+The classifier's real transcript ceiling is ~40-50K tokens (qwen-code's own
+`MAX_TRANSCRIPT_MESSAGES=40` × `MAX_HISTORICAL_ACTION_CHARS=4000`, per `fast-model-choice.md` §"what
+actually shipped") — nowhere near 500k/1M, so the classifier does **not** need to grow for this
+roadmap; it stays exactly as deployed today, on its own small allocation. Per-token cost is included
+here only because it feeds the "does the classifier's dedicated GPU need to change" question in §6.
+
+---
+
+## 4. Total VRAM budget: 500k × 2 parallel, and the 1M stretch
+
+Weights: `Qwen3.8-27B-UD-Q4_K_XL.gguf` is **17.6 GB**, confirmed directly from the
+[unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main)
+(already verified in `qwen3.8-27b-quant.md`).
+
+Per §2, both "500k × 2 parallel" and "1M stretch" require `--ctx-size 1000000` — same KV budget:
+
+| KV precision | KV cache | + weights (17.6 GB) | + est. compute-buffer/runtime overhead* | **Realistic total** |
+|---|---|---|---|---|
+| fp16 (default) | 61.0 GiB | 78.6 GiB | +3-6 GiB | **~82-85 GB** |
+| q8_0 (proven in this stack today) | 30.5 GiB | 48.1 GiB | +3-6 GiB | **~51-54 GB** |
+| q4_0 (proven in this stack today, on the classifier) | 15.3 GiB | 32.9 GiB | +3-6 GiB | **~36-39 GB** |
+
+\* *Estimate, not a cited figure* — llama.cpp's flash-attention compute buffer scales closer to
+linear than the unfused-attention path, per this repo's own measured note in `docker-compose.yml`'s
+`qwen-classifier` comment (unfused attention buffers ballooned unexpectedly at 65536 ctx; flash-attn
+fixed it). `--flash-attn on` is already the deployed default for the main model, so the linear-ish
+regime applies, but no primary source gives an exact formula for this buffer size at 1M context — the
++3-6 GiB band is this doc's estimate based on the ratio observed in that in-repo incident, not a
+llama.cpp-documented number. Budget for the high end of that range when sizing hardware.
+
+**Bottom line:** at `q4_0` KV (the most aggressive, already-proven-in-this-repo tier), 500k×2 /
+1M needs **~36-39 GB** total VRAM for the 27B model alone. That does not fit one 32GB card at any
+precision — it needs at least two 32GB-class cards, or one ≥40GB card. At `q8_0` (the precision this
+repo already runs in production for quality reasons), budget **~51-54 GB** — two 32GB cards (64GB
+pooled) clears this with room to spare; a single 48GB-class card would not.
+
+(§11 below extends this table to higher weight-quant tiers — Q6_K_XL, Q8_0, BF16 — for users who want
+better output quality than `Q4_K_XL`, and to a Flash-Next alternative architecture; see §11.6-§11.7.)
+
+---
+
+## 5. CPU/motherboard: which Threadripper generations give PCIe 4.0, and how many lanes for GPUs
+
+AMD's own product/chipset pages, cross-checked against the launch reviews that quote them directly:
+
+| Platform | Socket | PCIe generation | Total CPU-provided lanes |
+|---|---|---|---|
+| Threadripper 1000/2000 series ("1920X", "2950X", etc.) | **TR4** | **PCIe 3.0 only** | 60-64 |
+| Threadripper 3000 series (3960X/3970X/3990X) | **sTRX4** | **PCIe 4.0** | 64 |
+| Threadripper 7000 series (non-PRO) | sTR5 | PCIe 5.0 (48 lanes) + PCIe 4.0 (24-32 lanes) | ~72-80 |
+| Threadripper PRO 3000WX/5000WX | sWRX8 | PCIe 4.0 | **128** |
+| Threadripper PRO 7000WX | sTR5 (WRX90) | PCIe 5.0 (128 lanes) + a few PCIe 3.0 | **128** |
+
+Sources: [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html)
+(sWRX8/socket listing), corroborated by
+[Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch coverage](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2)
+(*"the 3rd Gen TR CPUs carry the same 64 PCIe lanes but double bandwidth by moving from Gen 3.0 to
+Gen 4.0"* — explicit confirmation TR4/1000-2000-series is PCIe 3.0 while sTRX4/3000-series is PCIe
+4.0), [PCWorld — Threadripper PRO launch](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html)
+(*"128 PCIe lanes"* for PRO).
+
+**This directly matters for the user's plan.** "An older Threadripper... for more PCIe lanes" is
+ambiguous between two real, very different chips:
+
+- **TR4 (1000/2000 series)** — cheapest used option, but **PCIe 3.0** — does not meet the user's own
+ stated PCIe 4.0 requirement, and PCIe 3.0 x8 per GPU roughly halves inter-GPU/host transfer
+ bandwidth (matters more for training/tensor-parallel than for llama.cpp's inference-time layer
+ splitting, but still a real downgrade vs. the R9700's native PCIe 5.0).
+- **sTRX4 (3000 series, non-PRO)** — the correct "older Threadripper with PCIe 4.0" target: 64 lanes,
+ 4-5 years old, real used-market availability, no PRO price premium.
+- **Threadripper PRO (3000WX/5000WX)** doubles the lane count to 128 but at meaningfully higher used
+ cost (workstation-tier, lower volume, sWRX8 boards are pricier than sTRX4/TRX40 boards) — worth it
+ only if 6 full-bandwidth (x16) GPU slots are actually needed; at x8-per-card (adequate for inference)
+ 64 lanes already covers 6 GPUs with lanes to spare for NVMe/chipset.
+
+**Lane budget for 6 GPUs on sTRX4 (64 lanes), estimated (no vendor spec gives a topology this
+specific — treat this bullet as an estimate):** typical sTRX4 boards reserve ~4 lanes for the
+chipset uplink and commonly wire 1-2 M.2 slots directly to the CPU (4 lanes each) — so realistic
+GPU-available lanes land around 44-52 of the 64, i.e. **6 GPUs at x8 electrical each (48 lanes) is
+plausible but board-model-dependent**; x16-each for 6 cards is not possible on 64 lanes regardless of
+board. x8 electrical is not a meaningful inference-speed penalty for llama.cpp (weights are loaded
+once; the ongoing per-token traffic across PCIe is small compared to compute), so this is an
+acceptable tradeoff, not a real bottleneck for this workload.
+
+---
+
+## 6. Power budget vs. the ~200W / one spare 8-pin headroom
+
+Official/vendor TDPs:
+
+| Card | TDP | Source |
+|---|---|---|
+| GTX 1080 (owned) | 180W, one 8-pin | [confirmed 180W, PCIe 3.0 x16, 1× 8-pin](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu) — spec matches NVIDIA's own launch figures reported across multiple outlets incl. Tom's Hardware |
+| RTX 2080 (owned) | 215W | Cross-checked across gpuzoo/cputronic/notebookcheck spec pages, consistent at 215W |
+| GT 710 (owned) | ~19W, **no external power connector** (slot power only) | [MSI/EVGA/Zotac GT 710 spec pages](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification) |
+| RTX 3060 12GB (candidate purchase) | 170W, one 8-pin | [NVIDIA-confirmed 170W TDP, one 8-pin connector](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/) |
+| RTX 3090 24GB (candidate purchase) | 350W, two 8-pin, [NVIDIA's own RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/) lists 350W and a 750W PSU minimum | NVIDIA official |
+| R9700 32GB (already deployed / "more of the same") | 300W (per this repo's `server-planing.md`, consistent with AMD's own R9700 product page framing it as a 300W-class card) | in-repo prior research |
+
+**Against the stated ~200W / one spare 8-pin budget:**
+
+- Adding **one RTX 3060 12GB** (170W, one 8-pin) is the *only* candidate in this list that fits the
+ stated headroom as-is — it uses the one spare connector and stays under 200W.
+- Adding the already-owned **GTX 1080** (180W) as the classifier's dedicated card also just barely
+ fits (180W ≤ 200W, one 8-pin) — this is a genuinely free option since the card is already owned and
+ its power draw is within budget, unlike every purchase candidate below.
+- Adding the already-owned **RTX 2080** (215W) **exceeds** the stated 200W headroom by 15W — technically
+ over budget on paper, though real-world draw is usually a bit under rated TDP; flag it as marginal,
+ not safely fitting.
+- Adding a **second R9700** (300W) or an **RTX 3090** (350W, needs two 8-pin — the user has only one
+ spare) both blow well past the current power budget on both watts and connector count.
+- **The GT 710 draws no meaningful power (~19W, no PCIe power connector at all)** — it is free from a
+ power-budget standpoint regardless of what else is added.
+
+**PSU upgrade trigger:** the very first stage that adds *any* GPU beyond a GTX 1080-class card (180W,
+one 8-pin) or an RTX 3060 12GB (170W, one 8-pin) exhausts the stated headroom. Any stage that reaches
+for a second 32GB-class card (R9700 or equivalent) or any 300W+ card **requires a PSU upgrade before
+that stage**, not after — see the roadmap table in §7 for exactly which stage that is.
+
+---
+
+## 7. Mixed-GPU feasibility: ROCm + CUDA, and is the GT 710 usable at all
+
+**ROCm and CUDA are different llama.cpp builds, but that's exactly the pattern already in this
+repo.** `ghcr.io/ggml-org/llama.cpp` publishes both `server-rocm` and `server-cuda` as separate,
+independently-built image tags (confirmed present on the [ggml-org container registry](https://github.com/orgs/ggml-org/packages/container/llama.cpp)
+and documented in [`docs/docker.md`](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md)
+— *"server-cuda: Same as `server` but compiled with CUDA support"*, *"server-rocm: Same as `server`
+but compiled with ROCm support"*). You cannot mix backends inside one process/container, but you
+**can** run one `server-rocm` container pinned to the R9700 and a separate `server-cuda` container
+pinned to an NVIDIA card, simultaneously, on the same host — this is architecturally identical to
+today's `llama-server` + `qwen-classifier` two-container split in `docker-compose.yml`, just with a
+different image tag for the NVIDIA-backed service and NVIDIA's container runtime (`nvidia-container-toolkit`
++ `--gpus` / device reservation, the CUDA-world equivalent of this repo's `/dev/kfd`+`/dev/dri`+
+numeric-GID ROCm pattern documented in
+[`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md)). None of that doc's ROCm-specific
+findings (the `GPU_MAX_HW_QUEUES=1` MES firmware workaround, the numeric-GID `group_add` fix) apply to
+an NVIDIA/CUDA container — those are ROCm-stack-specific bugs, not general multi-GPU-container issues.
+
+**Is this an implicit AMD→NVIDIA rebuild, or additive?** Worth surfacing explicitly since the two
+source plans conflict on this: `server-planing.md` is an AMD-only, ROCm-only 4×R9700 plan. This
+ticket's plan is **additive/mixed** — keep the R9700 running the main model under ROCm, and bolt on
+NVIDIA cards under CUDA for secondary roles (classifier, or a second inference GPU for the big model
+if going the "more of the same type" route means buying NVIDIA instead of more R9700s). Both are
+internally consistent, but they are different end-states — flag this choice back to the user rather
+than assuming one.
+
+**Splitting the *main* 27B model itself across mixed AMD+NVIDIA silicon in one process is not
+possible** — llama.cpp's multi-GPU tensor-split only works within a single backend build. To use
+both an R9700 and an NVIDIA card for the *same* model's layers, all the compute-hosting cards need to
+be the same backend (all-ROCm or all-CUDA) in that one process. This is why §5's roadmap treats "add
+GPU capacity to the main model" and "add a GPU for the classifier" as separable purchases with
+different backend constraints, not a single mixed pool.
+
+**Is the GT 710 usable for anything in this pipeline? No.** Reasoning:
+
+- 1GB VRAM cannot hold any meaningful fraction of either model's weights (17.6 GB / 2.4-4.3 GB) —
+ even a handful of transformer layers at Q4 quantization exceeds 1GB.
+- It's Kepler-generation silicon (192 CUDA cores, no tensor cores) — llama.cpp's CUDA backend
+ technically supports pre-Turing cards, but at this VRAM size there's nothing to usefully offload.
+- It draws power from the PCIe slot only, no external connector — genuinely free to keep installed.
+- **Plausible actual use: dedicate it as the box's display-output card**, so every compute-capable
+ GPU (R9700, and whichever NVIDIA cards get added) can be fully headless/compute-only with none of
+ their VRAM or a display output tied up driving a monitor — a real, if minor, use for it. This is
+ this doc's own inference from the spec facts above, not a claim found in any primary source.
+
+---
+
+## 8. GPU market pricing vs. the $20-30/GB assumption
+
+| Card | VRAM | Backend | Current used-market price (estimate — see caveat) | $/GB |
+|---|---|---|---|---|
+| RTX 3060 12GB | 12GB | CUDA | ~$240-300 used (eBay listings, [gpupoet.com tracker](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060): *"from $239"*, [eBay live listings](https://www.ebay.com/shop/rtx-3060-12gb) averaging ~$488 asking but with a $239 floor) | **~$20-25/GB** — matches the stated assumption |
+| RTX 3090 24GB | 24GB | CUDA | ~$1,010-1,050 used ([bestvaluegpu.com Sep 2026 tracker](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/), [xda-developers coverage](https://www.xda-developers.com/used-rtx-3090-still-best-for-local-ai-in-value/)) | **~$42-44/GB** — well above the stated assumption |
+| R9700 32GB ("more of the same type") | 32GB | ROCm | **New only — $1,299 MSRP**, street price $1,400-1,585 as of this research ([overclock3d](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/), [pricehistory.app tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm)) — too recent a release (2026) for a used market to exist yet | **~$41-50/GB, and not a used-market price at all** |
+
+**Caveat on all three price figures:** these are live marketplace asking-price snapshots pulled via
+web search on 2026-09-11, not sold-price data or a vendor spec sheet — treat as directional, not
+exact. eBay asking prices in particular run above realized sale prices.
+
+**Correction to the user's stated assumption:** $20-30/GB is a good estimate specifically for
+**last-generation mainstream used cards** (RTX 3060 12GB fits it almost exactly) but **not** for
+high-VRAM flagship cards like the RTX 3090 (~1.5-2× that rate) or for "more of the same type" R9700
+units, which aren't used-market at all yet and sit even higher per GB than the 3090. If the plan is
+"cheapest path to more VRAM," multiple RTX 3060 12GB cards (or similar mid-tier used cards) beat one
+RTX 3090 on $/GB, at the cost of needing more PCIe slots and more total wattage/connectors to reach
+the same aggregate VRAM — which is exactly the tradeoff the Threadripper lane-count plan in §5 is
+for.
+
+---
+
+## 9. Step-by-step roadmap
+
+All "resulting max context" figures assume `--parallel 2` and the KV precision stated; per §2, the
+`--ctx-size` value shown is the *total* (pre-division) value to pass to llama-server.
+
+| Stage | Hardware change | Est. cost | Backend | Usable VRAM (main-model pool) | Max context @ parallel=2 (`q4_0` KV) | PSU upgrade triggered? |
+|---|---|---|---|---|---|---|
+| **0 (current)** | 1× R9700 32GB, in production | $0 | ROCm | 32GB (shared with classifier) | ~131K/slot today at `q8_0` KV (262144 total ÷ 2) | No |
+| **1 — classifier isolation** | Move classifier onto the already-owned **GTX 1080** (180W, own container, `server-cuda`), freeing the R9700 entirely for the main model. Matches the existing dual-model pattern qwen-code's own docs describe (§10) and this repo's `qwen-classifier` service already implements, just on separate silicon instead of a shared card. | $0 (already owned) | ROCm (main) + CUDA (classifier) | R9700's full 32GB now available to the main model alone | ~262K/slot @ `q8_0` (unchanged ctx-size, no more classifier contention) | **No** — 180W GTX 1080 fits the stated ~200W/one-8-pin headroom |
+| **2 — second big-model GPU** | Add **one more 32GB-class card** for the main model. Cheapest correct-backend option: a second R9700 (~$1,300-1,585 new, ROCm, same backend as the first — required if tensor-splitting one model across two cards) | ~$1,300-1,585 | ROCm | 64GB pooled | `--ctx-size 500000 --parallel 1` fits at `q4_0` (~33GB) or `q8_0` (~48GB, tight but fits in 64GB) — **not yet 500k×2** | **Yes** — 300W card, no spare 8-pin left after stage 1 |
+| **3 — reach 500k × 2 / 1M stretch** | No further hardware if stage 2's 64GB pool is used with `--cache-type-k/v q4_0`: `--ctx-size 1000000 --parallel 2` needs ~33-39GB (§4), fits inside 64GB with real headroom for the compute buffer. If `q8_0` KV is required instead (this repo's current quality bar for the main model), the ~51-54GB need is tight-to-marginal on 64GB — a **third** 32GB card (~96GB pool) removes the risk. | $0 (reuses stage 2) or +$1,300-1,585 for a 3rd card if `q8_0` KV is required | ROCm | 64GB (q4_0 case) or 96GB (q8_0 case) | **500k×2 parallel achieved**, and the 1M stretch goal is the *same config* with `--parallel 1` instead of 2 (§2) | Already upgraded at stage 2 |
+| **4 — optional CPU/lane platform swap** | Only needed if the plan is to keep scaling past 2-3 big cards, or to add several small used cards (RTX 3060 12GB) for extra headroom/throughput rather than raw ctx-size. Swap to **non-PRO Threadripper 3000-series (sTRX4)** — 64 PCIe 4.0 lanes, ~x8-per-slot for up to 6 GPUs (§5). Threadripper PRO 3000WX/5000WX (128 lanes) only if x16-per-card matters or 6+ full-bandwidth slots are wanted. | Used sTRX4 CPU+board: roughly $400-800 combined on the used market (not independently priced in this pass — **estimate**, not cited) | n/a (platform only) | n/a | n/a | Independent of GPU wattage — driven by whatever GPU count/wattage stage 5+ adds |
+| **5+ — scale-out via small used cards** | Add RTX 3060 12GB units (~$20-25/GB, the assumption that actually holds, §8) instead of more 32GB flagship cards, once lane count (stage 4) supports it — useful for extra parallel slots / throughput beyond the 500k×2 target rather than for raising ctx-size further (500k×2/1M is already met at stage 3). | ~$240-300/card | CUDA (separate container per §6) | +12GB pooled per card, but on a *different backend* from the ROCm main model — usable for extra classifier/small-model capacity or a separate CUDA-backend llama-server instance, not as additional tensor-split VRAM for the ROCm main model | Unchanged for the main model; adds parallel capacity elsewhere | Yes, cumulative — each additional 170W card needs PSU headroom stage 2 already consumed |
+
+**Where the existing dual-model pattern sits in this roadmap:** it's stage 1, and it's free. The
+qwen-code docs pattern (main model + a small, always-resident, non-thinking fast/classifier model —
+see §10) is already implemented in this repo; the only roadmap-relevant change is *which GPU* the
+classifier sits on, moving it off the R9700 entirely onto an already-owned NVIDIA card frees the
+R9700's full 32GB for the 500k×2/1M push instead of splitting it with the classifier as happens
+today.
+
+### 9.1 Upgrade path, as diagrams
+
+Diagram form of the same §9 table and §11.9's dense-vs-Flash-Next call — nothing new is claimed here,
+this is a visual index back into the cited sections above.
+
+**Stage-by-stage hardware path** (PSU-upgrade triggers and target reached called out inline):
+
+```mermaid
+flowchart TD
+ S0["Stage 0 — today
1x R9700 32GB, ROCm
classifier shares the card
$0"]
+ S1["Stage 1 — classifier isolation
+ GTX 1080 (owned, 180W, CUDA)
R9700 freed for main model
$0 · PSU OK (180W fits ~200W headroom)"]
+ S2["Stage 2 — 2nd big-model GPU
+1x R9700 32GB (ROCm)
64GB pooled
~$1,300-1,585 · PSU UPGRADE REQUIRED (300W, no 8-pin left)"]
+ S3q4["Stage 3a — q4_0 KV
--ctx-size 1,000,000 --parallel 2
~33-39GB, fits in 64GB
$0 (reuses stage 2)"]
+ S3q8["Stage 3b — q8_0 KV (current prod quality)
~51-54GB, tight on 64GB
+1x R9700 -> 96GB removes risk
+~$1,300-1,585"]
+ TARGET(["500k x2 parallel reached
= 1M stretch goal, same VRAM
(--parallel 1 vs 2 is a config flag, §2)"])
+ S4["Stage 4 — platform swap (optional)
sTRX4 Threadripper 3000, 64 PCIe4 lanes
only needed past 2-3 big cards
~$400-800 (estimate, §9)"]
+ S5["Stage 5+ — scale out
+RTX 3060 12GB cards (CUDA, separate backend)
extra parallel/throughput, not more ctx-size
~$240-300/card · PSU upgrade each card"]
+
+ S0 --> S1 --> S2
+ S2 --> S3q4 --> TARGET
+ S2 --> S3q8 --> TARGET
+ TARGET -.->|"only if scaling past this"| S4 --> S5
+
+ style TARGET fill:#2e7d32,color:#fff,stroke:#1b5e20
+ style S2 fill:#8a5a00,color:#fff,stroke:#5c3d00
+ style S5 fill:#8a5a00,color:#fff,stroke:#5c3d00
+```
+
+**Model choice, and the one open question that could change it** (§11.9):
+
+```mermaid
+flowchart TD
+ Q{"Goal: 500k-1M ctx
within a $20-30/GB VRAM budget?"}
+ D["Dense Qwen3.8-27B
17.6-54.7GB weights (Q4_K_XL-BF16)
reaches 500k x2 on 2-3 cards,
500k x4 on 2-6 cards depending on quant
(§11.7 tables)"]
+ F{"Try --n-cpu-moe:
offload MoE experts to system RAM?
(untested for this model, §11.8)"}
+ FBAD["Flash-Next, all-GPU weights
111-354GB just for weights
needs 4-13 cards before any KV cost
NOT recommended at this budget (§11.9)"]
+ FGOOD["Flash-Next, experts in system RAM
GPU VRAM could shrink a lot
2.67x cheaper KV/token becomes relevant
UNVERIFIED — prototype on real server first"]
+ CAVEAT["+ real caveat either way:
PR #27742 flags unverified conv branch,
3% QSA divergence, prefill-pos-0-only PLE
(§11.1) — dense model carries no such flag"]
+
+ Q --> D
+ Q -->|"considering Flash-Next instead"| F
+ F -->|"works well"| FGOOD
+ F -->|"doesn't help / untested"| FBAD
+ FGOOD --> CAVEAT
+ FBAD --> CAVEAT
+
+ style D fill:#2e7d32,color:#fff,stroke:#1b5e20
+ style FBAD fill:#8a1c1c,color:#fff,stroke:#5c1212
+ style FGOOD fill:#8a5a00,color:#fff,stroke:#5c3d00
+```
+
+---
+
+## 10. Qwen-code's own docs on the fast-model/classifier pattern
+
+Fetched directly per the user's link:
+[qwenlm.github.io/qwen-code-docs/en/users/overview/](https://qwenlm.github.io/qwen-code-docs/en/users/overview/)
+— **the overview page itself does not describe the dual-model/classifier pattern**; it only covers
+single-model-provider setup (Alibaba ModelStudio / third-party / custom provider), one model at a
+time. The actual fast-model/classifier documentation lives on the **Auto Mode** page instead, which
+this repo's own `fast-model-choice.md` already fetched and cited in detail:
+[qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/) —
+summary (see `fast-model-choice.md` §1 for the full quote): a two-stage classifier gate, both stages
+using "your configured fast model (`/model --fast`)", Stage 1 a ~300ms `{shouldBlock}`-only check,
+Stage 2 a ~3-5s chain-of-thought review that only runs on a Stage-1 block. Nothing in either page
+gives a recommended *context size* or *model size* for the fast model beyond what's implied by that
+latency budget — this repo's own prior research (`fast-model-choice.md`) derived the actual context
+requirement from qwen-code's source code instead (`packages/core/src/permissions/classifier-transcript.ts`),
+since the docs pages don't state one. No new information changes that prior doc's conclusion; this
+section exists to confirm the overview page was checked directly as instructed and doesn't contradict
+or add to it.
+
+---
+
+## 11. Alternative: Qwen3.8-Flash-Next (MoE, hybrid attention)
+
+The user also wants to weigh switching (or adding) **Qwen3.8-Flash-Next** — a 125B-total/6B-active MoE
+with a hybrid recurrent-attention architecture — against staying on dense Qwen3.8-27B, and separately
+wants this section to cover **going up in weight quant** (Q4_K_XL → Q6_K_XL → Q8_0 → BF16/fp16) for
+*both* models, not just Q4. Feasibility first, since it gates everything else.
+
+### 11.1 Feasibility verdict: supported, but immature — read before trusting any number below
+
+Checked directly against the primary sources the task named:
+
+- **llama.cpp mainline support exists.** [PR #27742](https://github.com/ggml-org/llama.cpp/pull/27742)
+ ("model: add Qwen3.8-Flash-Next (qwen4exp)") was **merged into `master` on 2026-08-27** by ngxson.
+ It adds the full architecture: Gated DeltaNet layers (sigmoid-gated linear attention), QSA
+ ("Qwen Sparse Attention", operating at micro-block granularity), hyper-connections, and the PLE
+ n-gram embedding table. `llama.cpp`'s own docs list CPU/CUDA/Metal/ROCm as supported backends for
+ it — this is not a CUDA-only feature.
+- **This repo's pinned image is a floating tag, not a version pin.** `docker-compose.yml` runs
+ `ghcr.io/ggml-org/llama.cpp:server-rocm` with no date/digest suffix — a rolling "latest ROCm server
+ build" tag, not a release version. The merge is from 2026-08-27, and today is 2026-09-11 (~2 weeks
+ later), so a **fresh pull** of `server-rocm` should include it — but whatever image is already
+ cached/running on the R9700 box may predate the merge. **Action before touching this model on the
+ server: `docker compose pull llama-server` and check the startup log's build/commit banner is dated
+ on/after 2026-08-27**, not just "the tag says server-rocm."
+- **Real, primary-source-flagged immaturity — this is the part that should temper enthusiasm.** The
+ PR's own description/review discussion states: *"The conv branch itself is still numerically
+ unverified because the fixture zeroes its weights"*; QSA sparse attention *"diverges on 3 percent of
+ positions"* above its budget threshold; the PLE depthwise convolution *"is exact only for a prefill
+ that starts at position 0"* (i.e. correctness is not guaranteed once `--cache-reuse`/prompt-caching
+ is in play — a flag this repo already turns on for the dense model per the latest commit). None of
+ that is disqualifying, but it is a primary-source admission that this is a fresh, not-fully-verified
+ implementation, not a mature, widely-battle-tested one like the dense Qwen3.8-27B path.
+- **Multi-slot serving needs an explicit new flag.** The same PR states: *"`set_input_qsa` asserted
+ `n_stream == 1`, so llama-server could not serve this model with more than one slot unless `-kvu`
+ was passed."* Per the server README (§2), `--kv-unified`/`-kvu` defaults to enabled **only when slot
+ count is auto** (`-1`). This repo's compose file sets `--parallel ${LLAMA_PARALLEL:-2}` **explicitly**
+ (not auto) — so adopting Flash-Next with `--parallel` > 1 requires **adding `--kv-unified` (or
+ `-kvu`) to the launch flags**, a real deploy-time change, not something that "just works" by copying
+ today's flag set onto a new model file.
+
+**Verdict: yes, runnable** on this repo's backend (ROCm, mainline, no dev branch needed) as long as the
+image is pulled after 2026-08-27 and `-kvu` is added for multi-slot use — but treat it as
+**usable-with-caution**, not a drop-in swap, given the PR author's own unresolved-correctness notes.
+
+### 11.2 Architecture, verified against `config.json` directly
+
+Fetched from `Qwen/Qwen3.8-Flash-Next`'s `config.json` (unsloth's GGUF repo repackages the same base
+model): `num_hidden_layers=48`, `hidden_size=2560`, `num_attention_heads=24`, `num_key_value_heads=2`,
+`head_dim=256`, `max_position_embeddings=262144` (same native/extensible-to-1M framing as the dense
+model — same YaRN quality caveat from §3 applies here too, unverified past native length), `num_experts=512`,
+`num_experts_per_tok=10`, and the linear-attention head config: `linear_num_key_heads=16`,
+`linear_num_value_heads=48`, `linear_key_head_dim=128`, `linear_value_head_dim=128`.
+
+Layer pattern (confirmed both from the model card's own description and `config.json`'s
+`full_attention_interval=4`): every 4th layer is full/QSA attention, the other 3 are Gated DeltaNet —
+**12 of 48 layers grow a real KV cache; the other 36 have a fixed-size recurrent state that does not
+grow with context length.** (24% full-attention layers vs. the dense model's 16-of-64 = 25% — similar
+ratio, but the *absolute* per-layer KV cost differs because `num_key_value_heads` is 2 here vs. 4 on
+the dense model — see below.)
+
+### 11.3 Per-token growing-KV-cache cost
+
+```
+12 full-attention layers × 2 (K+V) × 2 kv_heads × 256 head_dim × 2 bytes (fp16) = 24 KiB/token
+```
+
+| Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 |
+|---|---|---|---|
+| 262,144 (native) | ~6.0 GiB | ~3.0 GiB | ~1.5 GiB |
+| 500,000 | ~11.4 GiB | ~5.7 GiB | ~2.9 GiB |
+| **1,000,000 (500k×2, or 1M stretch)** | **~22.9 GiB** | **~11.4 GiB** | **~5.7 GiB** |
+| **2,000,000 (500k×4)** | **~45.8 GiB** | **~22.9 GiB** | **~11.4 GiB** |
+
+`--cache-type-k/v` are the same generic llama.cpp KV-cache-quantization flags used elsewhere in this
+doc; nothing in the PR or the server README suggests they're handled differently for the 12
+full-attention layers of a hybrid model — they quantize the same growing K/V buffers as on a plain
+transformer. (No primary source explicitly confirms this for *this* architecture specifically — flagged
+as a reasonable extrapolation, not a directly-cited fact, same caveat class as this doc's other
+estimates.)
+
+### 11.4 Fixed (non-growing) recurrent state — Gated DeltaNet layers
+
+The 36 Gated DeltaNet layers each keep a fixed-size recurrent state (an outer-product-style
+key×value matrix per head) that does **not** scale with context length — only with slot/sequence
+count. Sized from `config.json`'s linear-attention head params:
+
+```
+36 layers × linear_num_value_heads(48) × linear_key_head_dim(128) × linear_value_head_dim(128) × 4 bytes (fp32 state)
+≈ 36 × 48 × 128 × 128 × 4 bytes ≈ 108 MiB per slot
+```
+
+This is **this doc's own derivation from the published head-dimension params, not a value pulled
+directly from llama.cpp source or docs** — the PR text confirms the state exists per-stream/per-slot
+but doesn't publish an exact byte formula, so treat the ~108 MiB/slot figure as an estimate, medium
+confidence. Even at 4 parallel slots that's under half a gigabyte — **negligible** next to both the
+growing KV cache (GBs) and the weights (tens to hundreds of GB) computed below. The headline
+implication holds regardless of the exact multiplier: Flash-Next's "big memory line item" is the MoE
+weights, not the attention state of any kind.
+
+### 11.5 Magnitude vs. the dense model — how much cheaper is KV, really
+
+At the same total ctx-size, Flash-Next's growing KV cache is **24 KiB/token vs. the dense model's
+64 KiB/token — 2.67× smaller**, i.e. Flash-Next's KV budget is **37.5%** of the dense model's at
+identical context length. This is a real, significant win *for the KV-cache line item specifically* —
+but see §11.9: it's a much smaller slice of a much bigger total, because the weights move the other
+way by a far larger factor.
+
+### 11.6 Weight sizes — verified from each unsloth GGUF repo's actual file listing
+
+Fetched directly from the HF file trees (not estimated from ratios), current as of this research pass:
+
+| Quant tier | Qwen3.8-27B (dense) | Qwen3.8-Flash-Next (MoE) |
+|---|---|---|
+| Q4_K_XL (`UD-Q4_K_XL`) | **17.6 GB** (existing baseline) | **111.4 GB** (4 parts: 10.9MB + 49.9GB + 49.4GB + 12.1GB) |
+| Q6_K_XL (`UD-Q6_K_XL`) | **25.3 GB** | **169 GB** (6 parts) |
+| Q8_0 | **29 GB** | **188 GB** (6 parts) |
+| BF16/fp16 | **54.67 GB** (50GB + 4.67GB, 2 parts) | **354 GB** (8 parts) |
+
+Sources: [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main)
+and its `BF16/` subfolder; [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main)
+and its `UD-Q4_K_XL/`, `UD-Q6_K_XL/`, `Q8_0/`, `BF16/` subfolders (per-file sizes summed). The
+preliminary Q8_0 figure floated before this research pass (~192GB) was slightly high — the real
+listing sums to **188 GB**; everything else in the preliminary list was accurate to within rounding.
+
+**The weight-quant axis and the KV-cache-quant axis are independent knobs.** Raising weight quality
+(Q4_K_XL → BF16) does not require raising `--cache-type-k/v` — the two flags are unrelated, and this
+repo already proves that pattern works (`q8_0` KV cache is deployed today against `Q4_K_XL` weights).
+A user chasing **maximum output quality** can run e.g. **BF16 weights + `q4_0` KV cache** — full-precision
+weights for quality, still-compressed KV for context budget — or any other combination in the tables
+below; nothing about picking a higher weight quant forces a matching KV precision.
+
+### 11.7 Total VRAM: does it fit, across quant tiers and both parallelism targets
+
+All totals = weights + growing KV cache + an estimated **+3-6 GB** compute-buffer/runtime overhead
+(same estimate band as §4, carried over — not re-derived for this architecture; flagged medium
+confidence there too). "Cards" = ceil(total ÷ 32GB), i.e. how many R9700-class 32GB cards it takes.
+
+#### Dense Qwen3.8-27B — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`)
+
+| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
+|---|---|---|---|
+| Q4_K_XL (17.6GB) | ~79-85GB (**3**) | ~48-54GB (**2**) | ~33-39GB (**2**) |
+| Q6_K_XL (25.3GB) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) | ~44-47GB (**2**) |
+| Q8_0 (29GB) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) | ~47-50GB (**2**) |
+| BF16 (54.67GB) | ~119-122GB (**4**) | ~88-91GB (**3**) | ~73-76GB (**3**) |
+
+#### Dense Qwen3.8-27B — 500k × 4 parallel (`--ctx-size 2,000,000`)
+
+| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
+|---|---|---|---|
+| Q4_K_XL (17.6GB) | ~143-146GB (**5**) | ~82-85GB (**3**) | ~51-54GB (**2**) |
+| Q6_K_XL (25.3GB) | ~150-153GB (**5**) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) |
+| Q8_0 (29GB) | ~154-157GB (**5**) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) |
+| BF16 (54.67GB) | ~180-183GB (**6**) | ~119-122GB (**4**) | ~88-91GB (**3**) |
+
+#### Qwen3.8-Flash-Next — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`)
+
+| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
+|---|---|---|---|
+| UD-Q4_K_XL (111.4GB) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) | ~120-123GB (**4**) |
+| UD-Q6_K_XL (169GB) | ~195-198GB (**7**) | ~183-186GB (**6**) | ~178-181GB (**6**) |
+| Q8_0 (188GB) | ~214-217GB (**7**) | ~202-205GB (**7**) | ~197-200GB (**7**) |
+| BF16 (354GB) | ~357-360GB (**12**) | ~357-360GB (**12**) | ~357-360GB (**12**) |
+
+#### Qwen3.8-Flash-Next — 500k × 4 parallel (`--ctx-size 2,000,000`)
+
+| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
+|---|---|---|---|
+| UD-Q4_K_XL (111.4GB) | ~160-163GB (**6**) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) |
+| UD-Q6_K_XL (169GB) | ~218-221GB (**7**) | ~195-198GB (**7**) | ~183-186GB (**6**) |
+| Q8_0 (188GB) | ~234-237GB (**8**) | ~211-214GB (**7**) | ~199-202GB (**7**) |
+| BF16 (354GB) | ~397-400GB (**13**) | ~377-380GB (**12**) | ~366-369GB (**12**) |
+
+(Flash-Next's KV precision barely moves the total at any weight quant above `UD-Q6_K_XL` — the weights
+so dominate the budget that KV quantization stops mattering for the "how many cards" question. This
+is the clearest signal in this whole section: for Flash-Next, the weight-quant choice is the entire
+hardware-sizing decision; for the dense model, KV precision still matters a lot.)
+
+### 11.8 CPU MoE-expert offload — the one lever that could change this calculus
+
+Flash-Next is a 512-expert/10-active-per-token MoE, and llama.cpp has a purpose-built flag for exactly
+this shape of model, confirmed directly from the server README: **`--n-cpu-moe`** — *"keep the Mixture
+of Experts (MoE) weights of the first N layers in the CPU"* — plus the more general
+**`--override-tensor`** (*"override tensor buffer type"*, pattern-matched by tensor name) that the same
+flag is built on top of. Both are generic, architecture-agnostic llama.cpp mechanisms (they match on
+tensor name patterns, not model type), so there's no reason to expect them not to apply to Flash-Next's
+MoE tensors specifically — but this pass found **no primary source that has actually tested
+`--n-cpu-moe` against this specific qwen4exp architecture**, so treat "it works here" as plausible,
+not confirmed.
+
+If it does work as expected, this changes the whole weight-VRAM picture in §11.7: the ~90-95% of
+Flash-Next's weight footprint that's MoE expert tensors could live in system RAM while attention
+projections, the shared/non-expert tensors, and the full KV cache stay on GPU — meaning a much smaller
+GPU-VRAM number than the "all weights on GPU" tables above, at the cost of PCIe/RAM-bandwidth-bound
+inference speed for whichever experts get selected per token (this repo has no benchmark of that
+tradeoff, and it's highly system-RAM-bandwidth-dependent, so no number is given here — flagged as an
+escape hatch worth prototyping directly on the server, not something this research values responsibly
+without a real test run).
+
+### 11.9 Net recommendation: dense Qwen3.8-27B vs. Flash-Next, for this user's stated goal
+
+**Net loss for this user's goal, as things stand — stay on dense Qwen3.8-27B.** Reasoning:
+
+- The user's target (500k×2 or 500k×4, on a $20-30/GB-VRAM budget, GPUs in 32GB increments) is a
+ **VRAM-budget-constrained** goal, and §11.7 shows Flash-Next's *weights alone* (111-354GB depending
+ on quant) dwarf the entire dense-model total-VRAM figure from §4/§11.7 (33-183GB depending on quant)
+ at every parallelism target. Flash-Next's much cheaper per-token KV cache (§11.5, real and verified)
+ is a rounding error next to that weight-size gap — the "2.67× cheaper KV" win doesn't come close to
+ offsetting a "6-20× larger weight footprint," so at $20-30/GB-VRAM the *dense* model reaches 500k×2
+ or 500k×4 for a fraction of the card count and dollar cost that Flash-Next needs even at its lowest
+ usable quant (`UD-Q4_K_XL`, 4-5 cards minimum) — before even factoring in §11.1's immaturity flags.
+- The one scenario that could flip this verdict is `--n-cpu-moe` actually working well for this
+ architecture (§11.8) — if most of those 111-354GB of expert weights can sit in system RAM at
+ acceptable throughput, Flash-Next's GPU-VRAM number could shrink dramatically and its real KV-cache
+ advantage would start to matter. That is untested here and shouldn't be assumed; it's the one
+ concrete next step worth trying on the actual server before ruling Flash-Next out permanently.
+- Independent of VRAM: §11.1's primary-source-flagged correctness caveats (unverified conv branch,
+ 3%-divergence QSA, prefill-position-0-only PLE exactness) are a real quality/stability risk on a
+ production coding-agent stack that dense Qwen3.8-27B simply doesn't carry, since it's been running
+ in this repo already.
+
+---
+
+## 12. 4-parallel × 500k scenario — all four combinations side by side
+
+Per §2's already-established, cited rule (`n_ctx_per_seq = n_ctx / n_parallel`,
+[ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681)), the same division
+applies at 4 slots: **500k tokens on each of 4 parallel slots requires `--ctx-size 2,000,000`** —
+double the 2-parallel target's `--ctx-size 1,000,000`, for the same reason 500k×2 needed double
+262,144. This isn't a new mechanism, just the same formula at `--parallel 4`.
+
+Full per-quant-tier tables for all four combinations are in §11.7 above (dense×2, dense×4, Flash-Next×2,
+Flash-Next×4 are each their own table there). Headline comparison at the KV precision already proven
+in production in this repo (`q8_0`) and each model's respective current/cheapest-usable weight quant:
+
+| Scenario | `--ctx-size` | Weight quant | q8_0-KV total VRAM | Cards (32GB) |
+|---|---|---|---|---|
+| Dense × 2 (or 1M stretch) | 1,000,000 | Q4_K_XL (17.6GB, current) | ~48-54GB | **2** |
+| Dense × 4 | 2,000,000 | Q4_K_XL (17.6GB, current) | ~82-85GB | **3** |
+| Flash-Next × 2 (or 1M stretch) | 1,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~126-129GB | **4**, edge |
+| Flash-Next × 4 | 2,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~137-140GB | **5** |
+
+**4-parallel × 500k reachability against this repo's existing roadmap stages (§9):**
+
+- **(a) Current 1×R9700 32GB:** none of the four combinations fit — not even dense×2 at any weight/KV
+ quant (§4's own conclusion, unchanged).
+- **(b) The 2-3×R9700 roadmap already proposed in §9 (64-96GB):** covers **dense×2 fully** (stage 3, as
+ already established) and **dense×4 at `q4_0` KV with Q4_K_XL or Q6_K_XL weights** (~51-62GB, fits in
+ 64-96GB) — but **not** dense×4 at higher weight quants (Q8_0/BF16 need 3-6 cards depending on KV
+ precision, per §11.7's dense×4 table) and **not any Flash-Next scenario** (minimum is 4 cards/128GB
+ even at the cheapest usable quant and tightest KV).
+- **(c) The full 4-6×R9700 stretch scenario** (`server-planing.md`'s original plan, 128-192GB pooled):
+ covers **dense×4 at every weight quant up to BF16** (worst case ~91GB at BF16+q4_0, well inside
+ 128GB) and **Flash-Next×2 at `UD-Q4_K_XL`** (126-140GB, fits a 5-card/160GB build, tight on a 4-card/
+ 128GB one) — but **not** Flash-Next×4 at any weight quant above `UD-Q4_K_XL`, and not Flash-Next at
+ `BF16` under any parallelism (needs 12-13 cards, an entirely different scale of build than anything
+ in this doc's roadmap).
+
+---
+
+## Sources
+
+- [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) — `--ctx-size`, `--parallel`, `--cache-type-k/v`, `--kv-unified`, `--cache-reuse`, `--n-cpu-moe`, `--override-tensor` flag definitions
+- [ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681) — real server log proving `n_ctx_per_seq = n_ctx / n_parallel`
+- [ggml-org/llama.cpp#27742](https://github.com/ggml-org/llama.cpp/pull/27742) — "model: add Qwen3.8-Flash-Next (qwen4exp)", merged 2026-08-27; architecture details, `n_stream == 1` / `-kvu` multi-slot requirement, and the conv-branch/QSA-divergence/PLE-prefill correctness caveats
+- [Qwen/Qwen3.8-27B config.json](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json)
+- [Qwen/Qwen3.8-Flash-Next config.json](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/raw/main/config.json)
+- [Qwen/Qwen3-4B-Instruct-2507 config.json](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json)
+- [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main) — 17.6GB (Q4_K_XL), 25.3GB (Q6_K_XL), 29GB (Q8_0), 54.67GB (BF16) weight sizes
+- [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main) — 111.4GB (UD-Q4_K_XL), 169GB (UD-Q6_K_XL), 188GB (Q8_0), 354GB (BF16) weight sizes, summed from each quant's per-file listing
+- [ggml-org/llama.cpp docs/docker.md](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md) — `server-cuda`/`server-rocm` separate image tags
+- [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html)
+- [Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2)
+- [PCWorld — Threadripper PRO launch, 128 PCIe lanes](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html)
+- [NVIDIA — GeForce RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/)
+- [Lowyat.net — RTX 3060 official 170W TDP](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/)
+- [BuildMyServer — GTX 1080 180W/PCIe3.0/1×8-pin spec listing](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu)
+- [MSI — GT 710 1GD5 LP spec page](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification)
+- [overclock3d — AMD Radeon AI PRO R9700 $1,299 MSRP](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/)
+- [pricehistory.app — R9700 street price tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm)
+- [bestvaluegpu.com — RTX 3090 used price tracker, Sep 2026](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/)
+- [gpupoet.com — RTX 3060 12GB used listings](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060)
+- [Qwen Code docs — overview](https://qwenlm.github.io/qwen-code-docs/en/users/overview/)
+- [Qwen Code docs — Auto Mode (fast-model pattern)](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/)
+- This repo: [`docker-compose.yml`](../../docker-compose.yml), [`.env.example`](../../.env.example), [`docs/server-planing.md`](../server-planing.md), [`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), [`qwen3.8-27b-tool-calling.md`](qwen3.8-27b-tool-calling.md), [`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md), [`fast-model-choice.md`](fast-model-choice.md)
+
+## Confidence/uncertainty summary
+
+- **High confidence:** the KV-cache-per-token formulas for both dense models and Flash-Next (computed
+ directly from each model's own `config.json`, same method this repo's prior research already used
+ and cross-checked); the `n_ctx_per_seq = n_ctx / n_parallel` division behavior (directly evidenced by
+ a real server log in a llama.cpp GitHub issue, and independently already documented in this repo's
+ own `.env.example`) — and confirmed to apply identically at `--parallel 4` since the mechanism is
+ parallel-count-agnostic; official TDP figures for GTX 1080, RTX 2080, RTX 3060, RTX 3090, GT 710
+ (each cross-checked against 2+ independent spec listings or the vendor's own product page); the
+ TR4-is-PCIe3/sTRX4-is-PCIe4 generational split (direct launch-coverage quote); the existence of
+ separate `server-cuda`/`server-rocm` llama.cpp image tags; Qwen3.8-Flash-Next's `config.json`
+ architecture params and PR #27742's merge date/status and its own stated correctness caveats and
+ `-kvu` multi-slot requirement (all directly quoted from the primary source); the weight file sizes
+ for both models at all four quant tiers (summed directly from each HF repo's real file listing, not
+ estimated).
+- **Medium confidence:** the compute-buffer/runtime-overhead estimate in §4/§11.7 (+3-6 GiB) —
+ extrapolated from one in-repo incident's before/after numbers, not a llama.cpp-documented formula,
+ and carried over to Flash-Next without re-derivation for its different architecture; the Gated
+ DeltaNet fixed recurrent-state size in §11.4 (~108 MiB/slot) — this doc's own derivation from the
+ published head-dimension config, not a value found in llama.cpp source or docs; whether
+ `--cache-type-k/v` quantization applies identically to Flash-Next's 12 full-attention layers as it
+ does to a plain transformer (reasonable extrapolation, not directly confirmed for this architecture);
+ whether `--n-cpu-moe`/`--override-tensor` actually work against Flash-Next's specific MoE tensor
+ layout (architecture-agnostic mechanism, but untested against this model by any primary source found);
+ real-world PCIe lane availability for 6 GPUs on a specific sTRX4 board (§5) — no single board's exact
+ lane map was fetched, this is a reasonable-but-unverified estimate from typical sTRX4 board behavior.
+- **Low confidence / explicitly estimated, not cited fact:** all used-GPU marketplace pricing (§8) —
+ live asking-price snapshots from a single search pass, not sold-price data; the used sTRX4
+ CPU+motherboard combo price in the roadmap's stage 4 (§9) — not researched at all in this pass,
+ flagged as a placeholder estimate; whether YaRN-scaled 500k/1M context actually holds output
+ quality for either Qwen3.8-27B or Qwen3.8-Flash-Next — no primary source (Qwen's own docs included)
+ publishes long-context quality benchmarks past the 262,144 native length for either model, so this is
+ a known-unknown carried forward from each model card's "YaRN-extensible" claim, not a verified
+ capability; whether the specific `ghcr.io/ggml-org/llama.cpp:server-rocm` image currently cached on
+ this repo's server actually postdates PR #27742's 2026-08-27 merge — not checked against the live
+ server in this pass, flagged as an action item in §11.1 rather than a confirmed fact.