Files
LLM-Server/docs/research/classifier-context-window-scaling.md

20 KiB
Raw Permalink Blame History

Does the qwen-classifier need to match the main model's context window, and would upgrading it to Qwen3-8B or Qwen3-Coder-30B-A3B-Instruct fit on the current R9700?

Date: 2026-09-15 Question raised: should qwen-classifier be resized to LLAMA_CTX_SIZE / LLAMA_PARALLEL (131072, matching the main model's per-slot context), should the classifier model itself move up to Qwen3-8B or Qwen3-Coder-30B-A3B-Instruct, and does that need a second GPU?

Answer: No, no, and not for this reason. qwen-code's own docs state no context-size requirement for fastModel at all — the "match the main model" premise doesn't come from any primary source. Separately, and independently of context size: neither Qwen3-8B nor Qwen3-Coder-30B-A3B-Instruct fits in the ~6.1GiB of VRAM actually free on the card today, at any context length, once weights alone are counted — this is a raw-VRAM problem, not a context-window problem, exactly matching the "qwen8b needs to offload more to the RAM" intuition in the request. A second GPU would solve the VRAM problem (and incidentally remove this repo's own GPU_MAX_HW_QUEUES=1 ROCm#5706 workaround from applying to this pair), but isn't deployed hardware today — it's a rack-build/acquisition question, not a config change.

1. Does qwen-code require the fast/classifier model to match the main model's context window?

No — checked against the user's own three linked pages, fetched directly:

  • fastModel settings docs: "Model used for generating prompt suggestions and speculative execution," configurable via inherit (main model), fast, a model ID, or authType:model-id; "Leave empty to use the main model." The docs recommend "a smaller/faster model (e.g., qwen3-coder-flash) reduces latency and cost" — no context-window size or capacity requirement is stated anywhere on this page. Source: qwen-code docs — Configuration / Settings, #fastmodel

  • Approval Mode / Auto Mode classifier docs: describes what the classifier evaluates (shell commands, network calls, out-of-workspace edits) and its allow/block behavior, but does not name a specific model or state any context-window requirement — the only operational note is that "when the classifier API is unreachable, the action is blocked rather than allowed." Source: qwen-code docs — Approval Mode, #4-auto-mode---classifier-driven-approval

  • Auto Mode "How it works": confirms the two-stage design (Stage 1: ~300ms, {shouldBlock} only; Stage 2: chain-of-thought reconsideration, only on a Stage-1 block) and what data reaches the classifier — user text, assistant tool-use calls, and tool-specific projections (truncated edit content, fetch URLs, shell command text). Tool results are explicitly never sent to the classifier. It "uses your configured fast model (/model --fast)," falling back to the main session model only if none is set. No statement anywhere requires or implies the fast model's context window match the main model's. Source: qwen-code docs — Auto Mode, #how-it-works

This confirms and sharpens what this repo's own fast-model-choice.md already found by reading qwen-code's source directly (packages/core/src/permissions/classifier-transcript.ts: MAX_TRANSCRIPT_MESSAGES=40, MAX_HISTORICAL_ACTION_CHARS=4000/message, worst case ~40-50K tokens, live-tested at 15,116 prompt tokens) — that doc already called the original 131072 in settings.json "copied from the main model's entry, not a real qwen-code requirement." The three docs pages fetched here add nothing that contradicts that: there is no primary-source basis for LLAMA_CTX_SIZE / LLAMA_PARALLEL symmetry between the two models. The current 65536 classifier ctx already carries ~1.5x margin over the real worst case.

2. VRAM math for Qwen3-8B and Qwen3-Coder-30B-A3B-Instruct as classifier candidates

Same method this repo already uses (qwen3.8-27b-quant.md, fast-model-choice.md §5): per-token KV cache = layers × 2(K+V) × kv_heads × head_dim × bytes, read directly from each model's own config.json.

Qwen3-8B

  • Architecture (Qwen/Qwen3-8B config.json): num_hidden_layers: 36, num_key_value_heads: 8, num_attention_heads: 32, head_dim: 128, hidden_size: 4096, max_position_embeddings: 40960, rope_scaling: null. Source: Qwen/Qwen3-8B config.json
  • Native context is 32,768 tokens, not the 262,144 the user's brief assumed (that number belongs to Qwen3-4B-Instruct-2507, a different, non-reasoning 2507-refresh model — Qwen3-8B is the earlier, thinking-capable Qwen3 architecture with a materially smaller native window). Extending past 32K needs YaRN:

    "Qwen3 natively supports context lengths of up to 32,768 tokens. For conversations where the total length (including both input and output) significantly exceeds this limit, we recommend using RoPE scaling techniques to handle long texts effectively." and llama.cpp-specific YaRN invocation is given explicitly: ./llama-cli ... -c 131072 --rope-scaling yarn --rope-scale 4 --yarn-orig-ctx 32768, with a documented caveat that "all the notable open-source frameworks implement static YaRN, which means the scaling factor remains constant regardless of input length, potentially impacting performance on shorter texts." Source: Qwen/Qwen3-8B-GGUF — Processing Long Texts

  • Weights (official Qwen quants, fetched from the GGUF repo file list): Q5_K_M = 5.85 GB, Q8_0 = 8.71 GB. Source: Qwen/Qwen3-8B-GGUF (Not independently verified: unsloth's equivalent UD-Q4_K_XL quant, which is what this repo's docker-compose.yml/.env.example actually download for every model deployed so far — the unsloth file size wasn't fetched, only the official Qwen quants above. Treat Q5_K_M/Q8_0 as a reasonable bound, not the exact file this repo would pull.)
  • Per-token KV cache: 36 × 2 × 8 × 128 × 2 bytes = 144 KiB/token fp16 — identical to Qwen3-4B-Instruct-2507's figure in fast-model-choice.md §5, since both share the same layers/kv_heads/head_dim triple.
Context KV (fp16) KV (q8_0) KV (q4_0, current classifier setting)
65,536 (current classifier ctx) 9.0 GiB 4.5 GiB 2.25 GiB
131,072 (user's proposed "match main model") 18.0 GiB 9.0 GiB 4.5 GiB

Weights + KV (q4_0, smallest realistic combo):

Context Q5_K_M weights + q4_0 KV Q8_0 weights + q4_0 KV
65,536 5.85 + 2.25 = 8.1 GB 8.71 + 2.25 = 10.96 GB
131,072 5.85 + 4.5 = 10.35 GB 8.71 + 4.5 = 13.21 GB

Qwen3-Coder-30B-A3B-Instruct

  • Architecture (fetched from the shared Qwen3-30B-A3B-family config.json): num_hidden_layers: 48, num_key_value_heads: 4, num_attention_heads: 32, head_dim: 128, hidden_size: 2048, MoE: num_experts: 128, num_experts_per_tok: 8 (8 of 128 experts active per token — confirms this is a sparse MoE model, not a dense one like the 27B or 8B candidates; the "active params" figure describes compute per token, not memory footprint — all 128 experts' weights still have to be resident wherever the model is loaded, GPU or RAM). Source: Qwen/Qwen3-30B-A3B-family config.json
  • UD-Q4_K_XL file size (the exact quant/quantizer this repo already standardizes on): 17.7 GB. 30.5B total / 3.3B activated parameters. Native context "262,144 natively... can be extended further using Yarn to reach 1M tokens." Source: unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
  • Per-token KV cache: 48 × 2 × 4 × 128 × 2 bytes = 96 KiB/token fp16 (smaller per-token than the 8B/4B candidates, since num_key_value_heads is 4 here vs. 8 — but this saving is irrelevant given the weights size below).
Context KV (fp16) KV (q8_0) KV (q4_0)
65,536 6.0 GiB 3.0 GiB 1.5 GiB
131,072 12.0 GiB 6.0 GiB 3.0 GiB

Weights + KV (q4_0):

Context Total
65,536 17.7 + 1.5 = 19.2 GB
131,072 17.7 + 3.0 = 20.7 GB

3. Does either candidate fit the ~6.1 GiB actually free on the card today?

No — neither does, at either context size, even at the smallest quant/KV-quant combination tested.

  • Qwen3-8B's cheapest realistic combination (Q5_K_M weights + q4_0 KV at the current 65536 ctx, not even the proposed 131072) is 8.1 GB — already ~2 GB over the measured 6.1 GiB free budget, before accounting for compute-buffer/batch overhead that fast-model-choice.md §"Implementation note" already found could add meaningfully on top of the naive weights+KV estimate (that's exactly why the 4B classifier ended up needing --flash-attn on and partial 28/36-layer offload instead of the originally-predicted comfortable full-GPU fit).
  • Qwen3-Coder-30B-A3B-Instruct isn't close at any setting tested — its weights alone (17.7 GB) are triple the entire free budget, and this doesn't change with context size since the weights term dominates.
  • This is a VRAM-capacity problem, not a context-window problem — directly confirming the "qwen8b needs to offload more to the RAM" intuition in the original request. Reducing context doesn't fix it; the weights don't fit regardless.

CPU/RAM offload mechanics: llama.cpp's --n-gpu-layers is documented as "max. number of layers to store in VRAM, either an exact number, 'auto', or 'all'" — the layers not selected are computed on CPU, with the model loaded via mmap by default (same mechanism this repo's own .env.example already documents for LLAMA_GPU_LAYERS: "if GPU+RAM ever can't hold the working set, the OS pages the rest in from disk automatically"). For the MoE Coder-30B-A3B model specifically, this repo's own .env.example already flags the more targeted alternative — --n-cpu-moe/--cpu-moe/--override-tensor "exps" — as the flags that "target Mixture-of-Experts models (e.g. Qwen3.8-2.4T-A95B)," i.e. exactly this model's architecture: these offload only expert-tensor weights to CPU while keeping attention/shared layers and KV cache on GPU, which is the mechanically correct lever for an MoE model, unlike the blunt --n-gpu-layers used for the dense 8B/27B/4B models. Neither llama.cpp's own README nor the Qwen model cards fetched here document a quantified performance cost for partial offload — no primary source gives a "N layers offloaded = X% slower" figure. Source: llama.cpp tools/server/README.md

What is directly measured, in this repo's own deployment history: CPU-only was tried first for the current, much smaller 4B classifier and rejected — "too slow in practice: real classification calls blew past OmniRoute's 60s timeout and retry-looped (504→499→504)" (docker-compose.yml's qwen-classifier comment block). An 8B dense model has roughly double the compute of the 4B model per token; a 30B-A3B model's routing overhead on CPU (choosing 8 of 128 experts per token, each a separate weight lookup) adds a different kind of cost that neither this repo nor the sources fetched here have measured. Given the classifier's Stage 1 has an explicit ~300ms latency budget (fast-model-choice.md §1, from qwen-code's own docs), and this repo already has one concrete data point that CPU offload breaks that budget at a smaller model size, extending either candidate onto significant CPU offload carries real, unquantified latency risk — the same failure mode already observed once, at a favorable (smaller) model size.

4. Does a second GPU solve this, and is one actually available?

Not today. docs/server-planing.md is a rack-build plan for "3-4x AMD Radeon AI PRO R9700 (32GB) GPUs" — a future-state document, not present inventory. Every GPU-facing comment in this repo's own docker-compose.yml/.env.example/rocm-gpu-pin-and-render-group.md consistently refers to "the single 32GB R9700" and measures the "~6.1GiB free" budget against one physical card holding both llama-server and qwen-classifier. Adding a second GPU is a hardware-acquisition and rack-build question — physically sourcing, installing, and power/PCIe-provisioning a card per server-planing.md's own build plan — not a docker-compose.yml/.env.example change.

If a second GPU were added, it would directly remove one already-documented risk for this specific pair: this repo's own rocm-gpu-pin-and-render-group.md traced the GPU-pinned-at-100%/ROCm#5706 bug to its precise trigger condition —

"The pin only appears with two concurrent HIP-context-holding processes on the same GPU... Root cause: an AMD MES (Micro Engine Scheduler) firmware bug triggered by HIP hardware-queue creation." Source: ROCm/ROCm#5706, via this repo's own rocm-gpu-pin-and-render-group.md

Since the confirmed trigger is two HIP contexts sharing one physical card, moving the classifier to its own, second GPU would put each service on a single-HIP-context card — the condition that trips the bug wouldn't exist for this pair anymore, and the GPU_MAX_HW_QUEUES=1 workaround currently applied to both services specifically because they share one card would no longer be load-bearing for this pair (it would still apply if any future third service shared a card with either model). This wasn't independently re-verified across two separate physical cards by any source fetched in this pass — it's a direct extrapolation from the confirmed root cause, same category of caveat that doc's own author already flagged for its within-one-card claim.

Bottom line / recommendation

  1. Don't apply LLAMA_CTX_SIZE / LLAMA_PARALLEL symmetry to the classifier. No qwen-code primary source states or implies the fast/classifier model needs to match the main model's context window. The real requirement (§1, already established in fast-model-choice.md) is ~40-50K tokens worst case; the current 65536 already has margin. Doubling to 131072 would only double VRAM spent on KV cache for a model that won't otherwise fit anyway (§2-3).
  2. Don't upgrade the classifier to Qwen3-8B or Qwen3-Coder-30B-A3B-Instruct on the current single-GPU setup. Neither fits the ~6.1GiB actually free, at any context size — this is a weights-size problem, not a context-window problem. Forcing it would mean either (a) shrinking the main 27B model's own VRAM footprint to make room (a real trade-off against the primary model, not evaluated here), or (b) CPU/partial offload, which this repo has direct, measured evidence already breaks the classifier's latency budget at a smaller model size than either candidate.
  3. A second GPU is the clean fix for VRAM contention and would also retire the ROCm#5706 workaround's relevance for this pair — but it isn't deployed hardware today. server-planing.md is a future build plan; this is an acquisition/rack-build decision, not something achievable via a config change right now.
  4. If the actual underlying motivation is classifier quality (not context capacity), that's a separate, legitimate question this doc doesn't answer — worth its own research pass rather than solving it via a bigger model that doesn't fit the hardware.

Sources

Confidence / uncertainty summary

  • High confidence: qwen-code's fastModel/Auto-Mode docs state no context-window requirement (direct quotes from all three linked pages); Qwen3-8B's native 32,768 context and YaRN caveat (direct model-card quote); Qwen3-Coder-30B-A3B-Instruct's MoE architecture and 17.7GB Q4_K_XL file size (direct from the quantizer's own repo page); the KV-cache-per-token math for both candidates (computed directly from each model's own config.json, same method already validated in this repo's prior research); the ROCm#5706 trigger condition being scoped to two HIP contexts on the same GPU (direct quote from this repo's own prior research, itself sourced from the upstream issue).
  • Medium confidence: the exact unsloth UD-Q4_K_XL-equivalent file size for Qwen3-8B — only the official Qwen quants (Q5_K_M/Q8_0) were fetched, not unsloth's own repo, so the real number this repo would actually download wasn't directly verified (bounded reasonably by the Q5_K_M figure, which is already the smallest realistic option and still doesn't fit). The claim that CPU/partial-offload latency risk scales unfavorably for larger/MoE models is a reasoned extrapolation from this repo's one measured data point (4B CPU-only rejected) plus general MoE-routing-overhead reasoning, not a directly measured benchmark for either candidate.
  • Low confidence / not independently verified: whether GPU_MAX_HW_QUEUES=1/ROCm#5706 genuinely has zero relevance across two separate physical GPUs (extrapolated from the confirmed same-GPU trigger condition, same caveat this repo's own prior research already flagged for its own claim); no primary source found that quantifies llama.cpp's actual inference-speed penalty for partial --n-gpu-layers or --n-cpu-moe offload in general — this is a documented gap in the sources checked, not a guessed number.