Compare commits
19
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a3f1099bfc | ||
|
|
ed90256f43 | ||
|
|
fa852c8fb4 | ||
|
|
2250e804db | ||
|
|
e31647812a | ||
|
|
930e407053 | ||
|
|
76043e2c6f | ||
|
|
1102273384 | ||
|
|
ba9ace6f71 | ||
|
|
8e2650807c | ||
|
|
4353e5c0e8 | ||
|
|
20cc0bcc70 | ||
|
|
c1e30ec9bb | ||
|
|
b3a64fe4b5 | ||
|
|
828bd4c046 | ||
|
|
16df051318 | ||
|
|
128503b68a | ||
|
|
df900404c0 | ||
|
|
f4729ba704 |
+22
-10
@@ -29,18 +29,25 @@ LLAMA_GPU_LAYERS=999
|
||||
# this size) — total ~25.6GB, ~6GB headroom, the same footprint the old
|
||||
# 131072 fp16 setting used. See docs/research/qwen3.8-27b-quant.md.
|
||||
LLAMA_CTX_SIZE=262144
|
||||
# Concurrent request slots. Was implicitly 4 (llama.cpp's compiled-in
|
||||
# default) with no flag set — under concurrent subagent fan-out, 4 requests
|
||||
# split the same GPU compute, so a large-context prefill can queue behind
|
||||
# others long enough to blow past OmniRoute's stream-idle timeout, which then
|
||||
# cancels the request (see issue-tracker notes on the timeout/cancel loop).
|
||||
# Dropped to 2 so each slot gets more compute and finishes prefill sooner;
|
||||
# raise back toward 4 if throughput (not latency) becomes the bottleneck
|
||||
# instead. Each slot gets LLAMA_CTX_SIZE / LLAMA_PARALLEL tokens of context —
|
||||
# real sessions have hit ~66K tokens, so don't drop LLAMA_CTX_SIZE without
|
||||
# checking that per-slot number stays comfortably above observed usage.
|
||||
# Concurrent request slots — the real hardware ceiling for this GPU, not a
|
||||
# tunable to raise for throughput (was implicitly 4, llama.cpp's compiled-in
|
||||
# default; dropped to 2 because more contended prefill was blowing requests
|
||||
# past OmniRoute's idle timeout — see OMNIROUTE_STREAM_IDLE_TIMEOUT_MS below).
|
||||
# The 3rd+ request now queues on llama.cpp itself instead — its own queue has
|
||||
# no timeout (tools/server/server-queue.cpp), it just waits for a slot — so
|
||||
# the timeout that matters moved to OmniRoute's per-connection
|
||||
# providerSpecificData.timeoutMs (dashboard/API only, not in this file; see
|
||||
# handoff notes in the issue tracker). Each slot gets LLAMA_CTX_SIZE /
|
||||
# LLAMA_PARALLEL tokens of context — real sessions have hit ~66K tokens, so
|
||||
# don't drop LLAMA_CTX_SIZE without checking that per-slot number stays
|
||||
# comfortably above observed usage.
|
||||
LLAMA_PARALLEL=2
|
||||
|
||||
# Dedicated GPU-resident backend for qwen-code's tool-call harmfulness
|
||||
# classifier (fastModel in ~/.qwen/settings.json) — see docker-compose.yml's
|
||||
# qwen-classifier service comment for the why and the VRAM/context math.
|
||||
LLAMA_CLASSIFIER_MODEL_FILE=Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf
|
||||
|
||||
# --- Lazytainer ---
|
||||
# Seconds of inactivity before llama-server is stopped. 900 = 15 min.
|
||||
LAZYTAINER_INACTIVE_TIMEOUT=900
|
||||
@@ -64,6 +71,11 @@ OMNIROUTE_DASHBOARD_PORT=20128
|
||||
# cancels it (which cancels the matching llama-server task too). 180s gives
|
||||
# contended prefill (see LLAMA_PARALLEL above) room to produce a first token.
|
||||
OMNIROUTE_STREAM_IDLE_TIMEOUT_MS=180000
|
||||
# Wait budget for the *first* SSE token specifically (distinct from the
|
||||
# inter-chunk timeout above) — see
|
||||
# docs/research/omniroute-non-ping-sse-stream-timeout.md. 30 min covers a
|
||||
# contended, large-context prefill even after retries eat into the budget.
|
||||
OMNIROUTE_REQUEST_TIMEOUT_MS=1800000
|
||||
# Random values, filled in automatically by ./scripts/update.sh — leave
|
||||
# blank. Bootstrap dashboard admin password (log in at the dashboard port,
|
||||
# change it there afterwards — this is only the first-boot value):
|
||||
|
||||
+1
-1
@@ -5,4 +5,4 @@
|
||||
data/
|
||||
.leankg/
|
||||
.cache/
|
||||
.qwen/temp
|
||||
.qwen/tmp
|
||||
@@ -30,7 +30,16 @@ services:
|
||||
--flash-attn on
|
||||
--cache-type-k q8_0
|
||||
--cache-type-v q8_0
|
||||
--cache-reuse 256
|
||||
--jinja
|
||||
# --cache-reuse 256: reuse cached KV for any matching prompt chunk of at
|
||||
# least 256 tokens (KV-shift, no reprocessing) instead of reprefilling
|
||||
# from scratch every request. Directly targets the actual root cause
|
||||
# behind the OmniRoute non-ping SSE timeout, not just the symptom — see
|
||||
# docs/research/omniroute-non-ping-sse-stream-timeout.md. Pairs with
|
||||
# OmniRoute's promptCacheAffinityEnabled (dashboard default), which keeps
|
||||
# a conversation's requests pinned to the same slot so there's a matching
|
||||
# prefix to reuse.
|
||||
# No published host port: llama-server is reached only via the omniroute
|
||||
# gateway on the ai-stack docker network now — see issue #15. Its
|
||||
# unauthenticated API no longer needs to be LAN-reachable directly.
|
||||
@@ -46,6 +55,82 @@ services:
|
||||
- "lazytainer.group.llamaserver.inactiveTimeout=${LAZYTAINER_INACTIVE_TIMEOUT:-900}"
|
||||
- "lazytainer.group.llamaserver.minPacketThreshold=2"
|
||||
|
||||
# Dedicated backend for qwen-code's tool-call harmfulness classifier
|
||||
# (fastModel in ~/.qwen/settings.json). Was aliased onto llama-server's own
|
||||
# 27B connection — every classification call then queued behind whatever
|
||||
# heavy generation was already running on that model's 2 GPU slots (issue
|
||||
# tracker: OmniRoute semaphore/pr-agent investigation).
|
||||
#
|
||||
# Tried CPU-only first (own process avoids the GPU queue entirely) — too
|
||||
# slow in practice: real classification calls blew past OmniRoute's 60s
|
||||
# timeout and retry-looped (504→499→504...). Moved to GPU instead.
|
||||
#
|
||||
# Context sizing: qwen-code's classifier transcript is hard-capped in its
|
||||
# own source (MAX_TRANSCRIPT_MESSAGES=40, MAX_HISTORICAL_ACTION_CHARS=4000
|
||||
# per message, packages/core/src/permissions/classifier-transcript.ts) —
|
||||
# worst case is ~40-50K tokens, nowhere near the 131072 originally set in
|
||||
# settings.json (that number was copied from the main model's entry, not
|
||||
# a real qwen-code requirement). 65536 ctx gives ~1.5x margin over that
|
||||
# worst case. Qwen3-4B-Instruct-2507 is still the model choice — smallest
|
||||
# Qwen3 with long native context (262144) without RoPE-scaling, in case
|
||||
# that margin ever needs to grow.
|
||||
#
|
||||
# VRAM: weights+KV math (~4.9GiB) predicted comfortable headroom in the
|
||||
# ~6.1GiB free on the R9700, but measured live it actually used ~5.85GiB —
|
||||
# left only ~700MB free, too tight. Dropping --batch-size/--ubatch-size
|
||||
# barely moved it (~768MB free) — wrong lever. Actual cause: llama-server
|
||||
# runs with --flash-attn on but this service was missing it — without
|
||||
# flash attention the unfused attention compute buffer at 65536 ctx is
|
||||
# much larger (roughly O(n^2) intermediate buffers vs flash-attn's fused,
|
||||
# near-linear workspace), dwarfing the naive weights+KV estimate. Added
|
||||
# --flash-attn on to match llama-server; re-verify with rocm-smi after
|
||||
# deploy before trusting any of these numbers again. GPU_MAX_HW_QUEUES=1
|
||||
# carried over from llama-server's comment above — same ROCm/ROCm#5706
|
||||
# clock-pinning bug applies now that two HIP contexts (this +
|
||||
# llama-server) share the card.
|
||||
#
|
||||
# --reasoning off is a no-cost safety net, not a confirmed-needed fix:
|
||||
# ggml-org/llama.cpp#20809 (closed) documents some server builds
|
||||
# misdetecting Qwen3-Instruct-2507 models as thinking models, routing
|
||||
# tool-call output into reasoning_content instead of tool_calls — exactly
|
||||
# the failure mode that ruled out the 27B model for this role in the
|
||||
# first place. Whether the current image build still has it was never
|
||||
# independently confirmed (see docs/research/fast-model-choice.md §4/§6).
|
||||
qwen-classifier:
|
||||
image: ghcr.io/ggml-org/llama.cpp:server-rocm
|
||||
container_name: qwen-classifier
|
||||
devices:
|
||||
- /dev/kfd
|
||||
- /dev/dri
|
||||
group_add:
|
||||
- "${HOST_VIDEO_GID:?run scripts/update.sh first to resolve this}"
|
||||
- "${HOST_RENDER_GID:?run scripts/update.sh first to resolve this}"
|
||||
security_opt:
|
||||
- seccomp=unconfined
|
||||
ipc: host
|
||||
environment:
|
||||
- GPU_MAX_HW_QUEUES=1
|
||||
volumes:
|
||||
- models:/models
|
||||
command: >
|
||||
-m /models/${LLAMA_CLASSIFIER_MODEL_FILE:-Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf}
|
||||
--host 0.0.0.0
|
||||
--port 8080
|
||||
--n-gpu-layers 28
|
||||
--ctx-size 65536
|
||||
--parallel 1
|
||||
--batch-size 512
|
||||
--ubatch-size 128
|
||||
--flash-attn on
|
||||
--reasoning off
|
||||
--cache-type-k q4_0
|
||||
--cache-type-v q4_0
|
||||
--jinja
|
||||
expose:
|
||||
- "8080"
|
||||
restart: unless-stopped
|
||||
networks: [ai-stack]
|
||||
|
||||
# ponytail: one-off downloader, not a standing service — run via
|
||||
# `docker compose --profile tools run --rm downloader`. Folded into
|
||||
# scripts/update.sh, which runs this every time; the `test -f` guard is
|
||||
@@ -67,6 +152,22 @@ services:
|
||||
curl -L --fail --create-dirs -o /models/${LLAMA_MODEL_FILE:-Qwen3.8-27B-UD-Q4_K_XL.gguf}
|
||||
https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/resolve/main/${LLAMA_MODEL_FILE:-Qwen3.8-27B-UD-Q4_K_XL.gguf}
|
||||
|
||||
# Same pattern as downloader above, separate service so this one small
|
||||
# file doesn't get re-checked/re-pulled by the big model's job.
|
||||
downloader-classifier:
|
||||
image: curlimages/curl:latest
|
||||
profiles: ["tools"]
|
||||
user: root
|
||||
volumes:
|
||||
- models:/models
|
||||
entrypoint: ["sh", "-c"]
|
||||
command:
|
||||
- >
|
||||
test -f /models/${LLAMA_CLASSIFIER_MODEL_FILE:-Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf} &&
|
||||
echo "already downloaded, skipping" ||
|
||||
curl -L --fail --create-dirs -o /models/${LLAMA_CLASSIFIER_MODEL_FILE:-Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf}
|
||||
https://huggingface.co/unsloth/Qwen3-4B-Instruct-2507-GGUF/resolve/main/${LLAMA_CLASSIFIER_MODEL_FILE:-Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf}
|
||||
|
||||
# Fetches the three Qwen-Image FP8 files ComfyUI needs (diffusion model,
|
||||
# text encoder, VAE) — same test -f guard pattern as downloader above.
|
||||
# See docs/research/image-generation-model-choice.md and issue #42.
|
||||
@@ -191,6 +292,16 @@ services:
|
||||
# LLAMA_PARALLEL above for the other half of this fix). Raised here so
|
||||
# it's tracked in git instead of a dashboard-only setting.
|
||||
- STREAM_IDLE_TIMEOUT_MS=${OMNIROUTE_STREAM_IDLE_TIMEOUT_MS:-180000}
|
||||
# Different timer than STREAM_IDLE_TIMEOUT_MS above — that one only
|
||||
# bounds gaps *between* SSE chunks once streaming has started.
|
||||
# REQUEST_TIMEOUT_MS bounds the wait for the *first* non-ping SSE
|
||||
# event, and it's what was still firing ("Stream produced no non-ping
|
||||
# SSE event within 95000ms") the morning after the timeout above was
|
||||
# raised — see docs/research/omniroute-non-ping-sse-stream-timeout.md.
|
||||
# Default 600000 (10 min) per OmniRoute's own docs, but the effective
|
||||
# deadline is remaining budget after retries/cooldowns eat into it, not
|
||||
# a flat timer, so raised well past the default for headroom.
|
||||
- REQUEST_TIMEOUT_MS=${OMNIROUTE_REQUEST_TIMEOUT_MS:-1800000}
|
||||
# Same reasoning as litellm's extra_hosts entry below — ai-stack's bridge
|
||||
# network can't resolve search.home on its own.
|
||||
extra_hosts:
|
||||
|
||||
@@ -31,4 +31,4 @@ Both serve the same underlying model — `Qwen3.8-27B-UD-Q4_K_XL.gguf`, register
|
||||
| [OpenCode](opencode.md) | OpenAI Chat Completions | `http://<ai-box>:${OMNIROUTE_PORT:-4000}/v1` | `opencode.json` provider block |
|
||||
| [Qwen Code](qwen-code.md) | OpenAI Chat Completions (2 models: chat + `fastModel`) | `http://<ai-box>:${OMNIROUTE_PORT:-4000}/v1` | `~/.qwen/settings.json` `modelProviders.openai` |
|
||||
|
||||
Further reading: `docs/research/qwen3.8-27b-tool-calling.md`, `docs/proxy-key-onboarding.md`.
|
||||
Further reading: `docs/research/qwen3.8-27b-tool-calling.md`, `docs/proxy-key-onboarding.md`, `docs/research/omniroute-account-semaphore-timeout.md` (a connection that can only handle a few concurrent requests — like `llama-server` or `qwen-classifier` — hits a hardcoded 30s reject once more requests queue up than its `maxConcurrent`, unless configured around it).
|
||||
|
||||
@@ -25,7 +25,7 @@ curl -fsSL https://opencode.ai/install | bash
|
||||
"models": {
|
||||
"qwen3.8-27b-local": {
|
||||
"name": "Qwen3.8-27B",
|
||||
"limit": { "context": 65536, "output": 8192 }
|
||||
"limit": { "context": 131072, "output": 8192 }
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -2,7 +2,18 @@
|
||||
|
||||
[← back to overview](index.md)
|
||||
|
||||
Qwen Code speaks plain **OpenAI Chat Completions**, and — unlike the other CLIs — needs *two* models: the main chat model, and a `fastModel` for Auto Mode's action classifier (a separate, always-resident, always-fast instance so classification doesn't queue behind chat prefill; see `docker-compose.yml`'s `llama-server-fast` service and `docs/research/fast-model-choice.md`). Both are registered as separate providers in OmniRoute but reachable through the same gateway URL. Config lives in `~/.qwen/settings.json`:
|
||||
Qwen Code speaks plain **OpenAI Chat Completions**, and — unlike the other CLIs — needs *two* models: the main chat model, and a `fastModel` for Auto Mode's action classifier. Both are registered as separate providers in OmniRoute but reachable through the same gateway URL.
|
||||
|
||||
## Why a second model exists
|
||||
|
||||
Auto Mode's action classifier (`permissions.autoMode`) is qwen-code's per-tool-call safety gate — it decides whether to auto-approve or block a shell command / tool call before it runs. It was originally aliased onto the main 27B model's own OmniRoute connection. That broke two ways in practice (see `docs/research/fast-model-choice.md` for the model research, and the issue-tracker history for the full incident):
|
||||
|
||||
- **Queued behind heavy work.** Every classification call competed for the main model's 2 GPU slots with whatever real generation was already running, so a classifier check could sit blocked for minutes.
|
||||
- **CPU-only was tried first and was too slow.** Isolating the classifier onto its own CPU-only llama.cpp instance avoided the GPU queue entirely, but real classification calls (which can carry a non-trivial conversation transcript, not just the bare tool call) blew past OmniRoute's request timeout and retry-looped.
|
||||
|
||||
The fix: a dedicated, GPU-resident `qwen-classifier` service (`docker-compose.yml`) running a small model (`Qwen3-4B-Instruct-2507`) on its own **partial** GPU offload — enough layers on the R9700 to be fast, sized to leave real VRAM headroom next to the 27B model rather than trusting a naive weights+KV estimate (see that service's comment block in `docker-compose.yml` for the actual measured numbers and the two wrong turns — batch-size tuning, then flash-attn — before partial offload turned out to be the real lever).
|
||||
|
||||
## `~/.qwen/settings.json`
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -16,12 +27,12 @@ Qwen Code speaks plain **OpenAI Chat Completions**, and — unlike the other CLI
|
||||
"generationConfig": { "contextWindowSize": 131072 }
|
||||
},
|
||||
{
|
||||
"id": "<fast-model-provider-id-in-omniroute>",
|
||||
"name": "qwen3.8-27b-classifier",
|
||||
"id": "<classifier-provider-id-in-omniroute>",
|
||||
"name": "qwen3-4b-classifier",
|
||||
"envKey": "OMNIROUTE_API_KEY",
|
||||
"baseUrl": "http://<ai-box>:${OMNIROUTE_PORT:-4000}/v1",
|
||||
"generationConfig": {
|
||||
"contextWindowSize": 8192,
|
||||
"contextWindowSize": 65536,
|
||||
"extra_body": { "chat_template_kwargs": { "enable_thinking": false } }
|
||||
}
|
||||
}
|
||||
@@ -32,15 +43,14 @@ Qwen Code speaks plain **OpenAI Chat Completions**, and — unlike the other CLI
|
||||
"name": "<main-model-provider-id-in-omniroute>",
|
||||
"baseUrl": "http://<ai-box>:${OMNIROUTE_PORT:-4000}/v1"
|
||||
},
|
||||
"fastModel": "<fast-model-provider-id-in-omniroute>"
|
||||
"fastModel": "<classifier-provider-id-in-omniroute>"
|
||||
}
|
||||
```
|
||||
|
||||
- `envKey` names the environment variable Qwen Code reads the virtual key from — set `OMNIROUTE_API_KEY=<qwen-code-cli virtual key>` before launching. Both providers can share one virtual key (as above); split it into two if you want separate usage tracking for chat vs. classifier calls.
|
||||
- **`contextWindowSize` is per-slot, not `LLAMA_CTX_SIZE` itself** — llama.cpp divides `--ctx-size` across `LLAMA_PARALLEL` concurrent slots, and each request only gets one slot's share (same correction applies to OpenCode's `limit.context`). Compute it per model from `.env`:
|
||||
- Main model: `LLAMA_CTX_SIZE / LLAMA_PARALLEL` = `262144 / 2` = **131072**.
|
||||
- Fast model: `LLAMA_FAST_CTX_SIZE / LLAMA_FAST_PARALLEL` = `8192 / 1` = **8192**. Undersizing this one specifically breaks Auto Mode ("Classifier stage 1 unavailable") once `hints.allow`/`softDeny`/`hardDeny` entries and recent-action history push a classifier call past it — see the `LLAMA_FAST_CTX_SIZE` comment in `.env.example` before raising it instead of `LLAMA_FAST_PARALLEL`.
|
||||
- `enable_thinking: false` on the fast model matters: the fast model file (`Qwen3-4B-Instruct-2507`) is already non-thinking, but this also suppresses `<think>` output on any fast-model swap that isn't, keeping classifier responses parseable.
|
||||
- **`contextWindowSize` for the main model is per-slot, not `LLAMA_CTX_SIZE` itself** — llama.cpp divides `--ctx-size` across `LLAMA_PARALLEL` concurrent slots, and each request only gets one slot's share (same correction applies to OpenCode's `limit.context`). Compute it from `.env`: `LLAMA_CTX_SIZE / LLAMA_PARALLEL` = `262144 / 2` = **131072**.
|
||||
- **The classifier's `contextWindowSize` (65536) is not per-slot math** — `qwen-classifier` runs `--parallel 1`, so its whole `--ctx-size` belongs to the one slot. 65536 isn't a guess either: qwen-code's own source hard-caps the classifier transcript (`MAX_TRANSCRIPT_MESSAGES=40`, `MAX_HISTORICAL_ACTION_CHARS=4000`/message in `packages/core/src/permissions/classifier-transcript.ts`) — worst case is ~40-50K tokens, so 65536 gives real margin without wasting VRAM the way the original 131072 (copied from the main model's entry, not an actual qwen-code requirement) would have.
|
||||
- `enable_thinking: false` on the classifier matters for parseability, though `Qwen3-4B-Instruct-2507` is already architecturally non-thinking (see `fast-model-choice.md` §3) — this is belt-and-suspenders for any future fast-model swap that isn't.
|
||||
- Qwen Code also recognizes `advisorModel`, `visionModel`, `compactionModel`, `imageModel` for other model roles — none are wired up in this stack; only `fastModel` is required.
|
||||
|
||||
## Web search via OmniRoute
|
||||
@@ -96,9 +106,11 @@ Register it in `~/.qwen/settings.json`:
|
||||
|
||||
It reuses the same `OMNIROUTE_API_KEY` env var as the model providers above — the virtual key needs search permission in OmniRoute, not just chat-completions.
|
||||
|
||||
**Non-interactive mode (`qwen -p ...`) needs this tool explicitly allow-listed.** MCP tools require interactive confirmation by default; `--approval-mode auto` alone doesn't bypass that for a non-interactive run — pass `--allowed-tools mcp__omniroute-search__search` (or `-y` for full YOLO) alongside `-p`, or the search call never reaches the classifier at all and silently no-ops. Confirmed live: without the allow-list, only the tool calls the CLI's non-interactive gate lets through end up as classifier requests.
|
||||
|
||||
## Auto Mode tuning
|
||||
|
||||
Auto Mode's action classifier calls the fast model above — its own request can queue behind other stack traffic before the fast llama-server instance is warm, so the default classifier timeout is worth raising. And since this stack is a single trusted local proxy, it's reasonable to pre-approve requests to it rather than confirm every call:
|
||||
Auto Mode's action classifier calls the fast model above. Even on the dedicated GPU-resident instance, give it real timeout headroom rather than trusting OmniRoute's default — and since this stack is a single trusted local proxy, it's reasonable to pre-approve requests to it rather than confirm every call:
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -111,6 +123,8 @@ Auto Mode's action classifier calls the fast model above — its own request can
|
||||
}
|
||||
```
|
||||
|
||||
`hints.allow` entries are free-text descriptions the classifier matches against, not exact strings — capped at 150 entries/200 chars each (see the `LLAMA_FAST_CTX_SIZE` note above for why that ceiling matters).
|
||||
`hints.allow` entries are free-text descriptions the classifier matches against, not exact strings — capped at 150 entries/200 chars each.
|
||||
|
||||
Also set a generous per-connection timeout on the classifier's own OmniRoute provider connection (`providerSpecificData.timeoutMs`, dashboard or `PATCH /api/providers/{id}` — not a `.env` value, see `docs/network-access.md` for reaching the dashboard API). 120000ms is comfortable for the current GPU-resident setup (real measured latency: well under a second for a short check, low seconds for the largest realistic transcript) — this isn't the 20-minute figure the main 27B connection needs, since the classifier isn't competing for a contended GPU slot the way the main model can.
|
||||
|
||||
Everything else in `~/.qwen/settings.json` (`hooks`, `security.auth`'s underlying tooling, editor prefs) is per-machine, not part of pointing at this stack — don't copy it wholesale between machines.
|
||||
|
||||
@@ -0,0 +1,737 @@
|
||||
# Research: hardware roadmap to 500k-token context × 2 parallel agents (1M stretch)
|
||||
|
||||
**Date:** 2026-09-11
|
||||
|
||||
**Question:** What VRAM does 500k-token context × 2 parallel llama-server slots (and a 1M-token
|
||||
stretch goal) actually cost for the Qwen3 family, and what hardware roadmap gets there from the
|
||||
current single-R9700 setup — given the user's stated plan to add an older (PCIe 4.0) Threadripper
|
||||
for lane count, reuse existing RAM/PSU (~200W headroom / one spare 8-pin), mix in already-owned
|
||||
NVIDIA cards (GTX 1080 8GB, RTX 2080 8GB, GT 710 1GB) for the classifier role, and price used GPUs
|
||||
at roughly $20-30/GB VRAM?
|
||||
|
||||
**Answer, short version:** The two goals ("500k × 2 parallel" and "1M stretch") turn out to need
|
||||
**the same total VRAM budget** — because of how llama-server's `--ctx-size` and `--parallel` interact
|
||||
(§2), 500k × 2 slots and a single 1M-token slot both require setting `--ctx-size 1000000`. At
|
||||
`q4_0`-quantized KV cache that's **~33 GB** (weights + KV) for Qwen3.8-27B, at `q8_0` it's **~48 GB**,
|
||||
at fp16 it's **~79 GB** — before compute-buffer overhead. That does not fit on the current single
|
||||
32GB R9700 at any KV precision, and comfortably fits on two 32GB-class cards only at `q8_0`/`q4_0`.
|
||||
The user's $20-30/GB pricing intuition holds for last-gen used consumer cards (RTX 3060 12GB) but
|
||||
**not** for RTX 3090 24GB (~$44/GB currently) or a second R9700 (~$41/GB, new — no used market yet
|
||||
for a card released mid-2026). The stated Threadripper plan needs to specifically target the
|
||||
**non-PRO Threadripper 3000 series on sTRX4** (64 lanes, PCIe 4.0) — older Threadripper on the
|
||||
original TR4 socket (1000/2000 series) is PCIe 3.0 only, which doesn't match the user's own PCIe 4.0
|
||||
requirement. The power budget (~200W / one spare 8-pin) is exhausted by a *single* mid-tier used GPU
|
||||
addition — a PSU upgrade is not optional past the very first stage. See §7 for the roadmap.
|
||||
|
||||
---
|
||||
|
||||
## 1. Current state (from this repo)
|
||||
|
||||
From `docker-compose.yml` and `.env.example` at the repo root:
|
||||
|
||||
- **Main model:** `Qwen3.8-27B-UD-Q4_K_XL.gguf` (17.6 GB weights), `--ctx-size 262144`,
|
||||
`--parallel 2`, `--flash-attn on`, `--cache-type-k q8_0 --cache-type-v q8_0`, `--n-gpu-layers 999`,
|
||||
on one AMD Radeon AI PRO R9700 (32GB, ROCm/HIP, `gfx1201`).
|
||||
- **Classifier ("fast") model:** `Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`, `--ctx-size 65536`,
|
||||
`--parallel 1`, `--n-gpu-layers 28` (partial offload), `--cache-type-k/v q4_0`, its own container on
|
||||
the *same* R9700, sharing VRAM with the main model — see
|
||||
[`docker-compose.yml`](../../docker-compose.yml) lines ~58-99 and
|
||||
[`fast-model-choice.md`](fast-model-choice.md).
|
||||
- `.env.example` already documents the exact fact this research turns on: *"Each slot gets
|
||||
`LLAMA_CTX_SIZE / LLAMA_PARALLEL` tokens of context"* — i.e. today's 262144 ctx-size ÷ 2 parallel
|
||||
slots means each real request only gets **~131K tokens**, not the full 262144, confirmed in-repo
|
||||
before any external source was checked.
|
||||
- Prior research already worked out the KV-cache formula for this exact model
|
||||
([`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md)) — this doc reuses and extends that math for the
|
||||
500k/1M targets rather than re-deriving it.
|
||||
|
||||
`docs/server-planing.md` describes a **different, earlier plan**: a 4× AMD Radeon AI PRO R9700 rig
|
||||
on a Gigabyte MZ32-AR0 (single-socket SP3/EPYC, 128 PCIe 4.0 lanes), fully AMD/ROCm. The user's plan
|
||||
in this ticket is not that — it pivots toward an older **Threadripper** (SP3's sibling desktop-HEDT
|
||||
socket family, not SP3 itself) and explicitly wants to mix in already-owned **NVIDIA** cards. These
|
||||
two plans are **not the same build** and, per §6, ROCm and CUDA cards cannot share one llama.cpp
|
||||
process — they can only coexist as separate containers on separate cards. Treat `server-planing.md`
|
||||
as superseded context, not the active plan, unless the user says otherwise.
|
||||
|
||||
---
|
||||
|
||||
## 2. llama-server parallelism: does each slot get its own full `--ctx-size`, or is it divided?
|
||||
|
||||
**Divided.** This is the single fact that changes the whole budget by 2×, confirmed from three
|
||||
independent primary sources:
|
||||
|
||||
1. **This repo's own `.env.example`** (quoted above) already documents it for the current deployment.
|
||||
2. **llama.cpp's own server README**, fetched directly
|
||||
([`tools/server/README.md`](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)):
|
||||
`--ctx-size (-c)`: *"size of the prompt context (default: 0, 0 = loaded from model)"*;
|
||||
`--parallel (-np)`: *"number of server slots (default: -1, -1 = auto)"* — the docs list these as
|
||||
independent flags, but don't spell out the division themselves.
|
||||
3. **A real user's server log**, quoted verbatim in
|
||||
[ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681), is the actual proof:
|
||||
running with `--ctx-size 327680 --parallel 6` produces `n_ctx = 327680`,
|
||||
`n_ctx_per_seq = 54613` — i.e. `327680 / 6 ≈ 54613`. The reporter explicitly asked for a
|
||||
`--ctx-size-per-seq`-style flag to *avoid* this division; no such flag exists as of the fetch date.
|
||||
|
||||
Practical consequence: **to get 500,000 usable tokens on each of 2 parallel slots, `--ctx-size` must
|
||||
be set to 1,000,000, not 500,000.** The KV cache is sized off the *total* `--ctx-size`
|
||||
(`--kv-unified`, on by default when slots are auto per the README's `-kvu` entry, uses one shared
|
||||
pool sized to the full `n_ctx`) — so the VRAM cost of "500k × 2 parallel" and "one 1M-token slot"
|
||||
is **identical**: both require `--ctx-size 1000000`. This is a genuinely useful finding for the
|
||||
roadmap — reaching the 500k×2 target and the 1M stretch goal cost the same VRAM; the only difference
|
||||
is `--parallel 1` vs `--parallel 2` at deploy time, a config change with zero extra hardware cost.
|
||||
|
||||
`--cache-type-k` / `--cache-type-v` accept `f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1`
|
||||
(default `f16`), per the same README fetch. The repo already uses `q8_0` on the main model and `q4_0`
|
||||
on the classifier, so both quantization tiers used in the math below are already-proven-working
|
||||
configurations in this stack, not hypothetical flags.
|
||||
|
||||
---
|
||||
|
||||
## 3. KV-cache math per model
|
||||
|
||||
### Qwen3.8-27B (hybrid Gated-DeltaNet / attention)
|
||||
|
||||
Reusing the architecture params already pulled from
|
||||
[`Qwen/Qwen3.8-27B/config.json`](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json) in
|
||||
[`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), re-verified directly for this doc: `num_hidden_layers=64`,
|
||||
`full_attention_interval=4` → **16 of 64 layers are standard KV-caching attention**, the other 48 are
|
||||
Gated DeltaNet linear-attention layers with a small, context-length-*independent* recurrent state
|
||||
(tens of MB total, negligible next to the attention KV cache — ignored below).
|
||||
`num_key_value_heads=4` (GQA), `head_dim=256`. Native context `max_position_embeddings=262144`
|
||||
(YaRN-extensible to 1M per the model card — **both the 500k and 1M targets exceed native context and
|
||||
require RoPE/YaRN scaling**, which is a real quality caveat, not just a memory one — Qwen has not
|
||||
published independent long-context quality benchmarks past native length that this research found).
|
||||
|
||||
Per-token KV cache, fp16, both K and V, across the 16 full-attention layers:
|
||||
|
||||
```
|
||||
16 layers × 2 (K+V) × 4 kv_heads × 256 head_dim × 2 bytes = 64 KiB/token
|
||||
```
|
||||
|
||||
| Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 |
|
||||
|---|---|---|---|
|
||||
| 262,144 (current) | ~16.0 GiB | ~8.0 GiB | ~4.0 GiB |
|
||||
| 500,000 | ~30.5 GiB | ~15.3 GiB | ~7.6 GiB |
|
||||
| **1,000,000 (500k×2, or 1M stretch)** | **~61.0 GiB** | **~30.5 GiB** | **~15.3 GiB** |
|
||||
|
||||
(`q8_0` is 8-bit vs. fp16's 16-bit → exactly half; `q4_0` is 4-bit → exactly quarter, per llama.cpp's
|
||||
own cache-type byte widths.)
|
||||
|
||||
### Qwen3-4B-Instruct-2507 (plain GQA transformer, classifier role)
|
||||
|
||||
From [`Qwen/Qwen3-4B-Instruct-2507/config.json`](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json)
|
||||
(already pulled in [`fast-model-choice.md`](fast-model-choice.md)): `num_hidden_layers=36` — every
|
||||
layer is standard attention here (no hybrid split), `num_key_value_heads=8`, `head_dim=128`.
|
||||
|
||||
```
|
||||
36 layers × 2 (K+V) × 8 kv_heads × 128 head_dim × 2 bytes = 144 KiB/token
|
||||
```
|
||||
|
||||
The classifier's real transcript ceiling is ~40-50K tokens (qwen-code's own
|
||||
`MAX_TRANSCRIPT_MESSAGES=40` × `MAX_HISTORICAL_ACTION_CHARS=4000`, per `fast-model-choice.md` §"what
|
||||
actually shipped") — nowhere near 500k/1M, so the classifier does **not** need to grow for this
|
||||
roadmap; it stays exactly as deployed today, on its own small allocation. Per-token cost is included
|
||||
here only because it feeds the "does the classifier's dedicated GPU need to change" question in §6.
|
||||
|
||||
---
|
||||
|
||||
## 4. Total VRAM budget: 500k × 2 parallel, and the 1M stretch
|
||||
|
||||
Weights: `Qwen3.8-27B-UD-Q4_K_XL.gguf` is **17.6 GB**, confirmed directly from the
|
||||
[unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main)
|
||||
(already verified in `qwen3.8-27b-quant.md`).
|
||||
|
||||
Per §2, both "500k × 2 parallel" and "1M stretch" require `--ctx-size 1000000` — same KV budget:
|
||||
|
||||
| KV precision | KV cache | + weights (17.6 GB) | + est. compute-buffer/runtime overhead* | **Realistic total** |
|
||||
|---|---|---|---|---|
|
||||
| fp16 (default) | 61.0 GiB | 78.6 GiB | +3-6 GiB | **~82-85 GB** |
|
||||
| q8_0 (proven in this stack today) | 30.5 GiB | 48.1 GiB | +3-6 GiB | **~51-54 GB** |
|
||||
| q4_0 (proven in this stack today, on the classifier) | 15.3 GiB | 32.9 GiB | +3-6 GiB | **~36-39 GB** |
|
||||
|
||||
\* *Estimate, not a cited figure* — llama.cpp's flash-attention compute buffer scales closer to
|
||||
linear than the unfused-attention path, per this repo's own measured note in `docker-compose.yml`'s
|
||||
`qwen-classifier` comment (unfused attention buffers ballooned unexpectedly at 65536 ctx; flash-attn
|
||||
fixed it). `--flash-attn on` is already the deployed default for the main model, so the linear-ish
|
||||
regime applies, but no primary source gives an exact formula for this buffer size at 1M context — the
|
||||
+3-6 GiB band is this doc's estimate based on the ratio observed in that in-repo incident, not a
|
||||
llama.cpp-documented number. Budget for the high end of that range when sizing hardware.
|
||||
|
||||
**Bottom line:** at `q4_0` KV (the most aggressive, already-proven-in-this-repo tier), 500k×2 /
|
||||
1M needs **~36-39 GB** total VRAM for the 27B model alone. That does not fit one 32GB card at any
|
||||
precision — it needs at least two 32GB-class cards, or one ≥40GB card. At `q8_0` (the precision this
|
||||
repo already runs in production for quality reasons), budget **~51-54 GB** — two 32GB cards (64GB
|
||||
pooled) clears this with room to spare; a single 48GB-class card would not.
|
||||
|
||||
(§11 below extends this table to higher weight-quant tiers — Q6_K_XL, Q8_0, BF16 — for users who want
|
||||
better output quality than `Q4_K_XL`, and to a Flash-Next alternative architecture; see §11.6-§11.7.)
|
||||
|
||||
---
|
||||
|
||||
## 5. CPU/motherboard: which Threadripper generations give PCIe 4.0, and how many lanes for GPUs
|
||||
|
||||
AMD's own product/chipset pages, cross-checked against the launch reviews that quote them directly:
|
||||
|
||||
| Platform | Socket | PCIe generation | Total CPU-provided lanes |
|
||||
|---|---|---|---|
|
||||
| Threadripper 1000/2000 series ("1920X", "2950X", etc.) | **TR4** | **PCIe 3.0 only** | 60-64 |
|
||||
| Threadripper 3000 series (3960X/3970X/3990X) | **sTRX4** | **PCIe 4.0** | 64 |
|
||||
| Threadripper 7000 series (non-PRO) | sTR5 | PCIe 5.0 (48 lanes) + PCIe 4.0 (24-32 lanes) | ~72-80 |
|
||||
| Threadripper PRO 3000WX/5000WX | sWRX8 | PCIe 4.0 | **128** |
|
||||
| Threadripper PRO 7000WX | sTR5 (WRX90) | PCIe 5.0 (128 lanes) + a few PCIe 3.0 | **128** |
|
||||
|
||||
Sources: [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html)
|
||||
(sWRX8/socket listing), corroborated by
|
||||
[Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch coverage](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2)
|
||||
(*"the 3rd Gen TR CPUs carry the same 64 PCIe lanes but double bandwidth by moving from Gen 3.0 to
|
||||
Gen 4.0"* — explicit confirmation TR4/1000-2000-series is PCIe 3.0 while sTRX4/3000-series is PCIe
|
||||
4.0), [PCWorld — Threadripper PRO launch](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html)
|
||||
(*"128 PCIe lanes"* for PRO).
|
||||
|
||||
**This directly matters for the user's plan.** "An older Threadripper... for more PCIe lanes" is
|
||||
ambiguous between two real, very different chips:
|
||||
|
||||
- **TR4 (1000/2000 series)** — cheapest used option, but **PCIe 3.0** — does not meet the user's own
|
||||
stated PCIe 4.0 requirement, and PCIe 3.0 x8 per GPU roughly halves inter-GPU/host transfer
|
||||
bandwidth (matters more for training/tensor-parallel than for llama.cpp's inference-time layer
|
||||
splitting, but still a real downgrade vs. the R9700's native PCIe 5.0).
|
||||
- **sTRX4 (3000 series, non-PRO)** — the correct "older Threadripper with PCIe 4.0" target: 64 lanes,
|
||||
4-5 years old, real used-market availability, no PRO price premium.
|
||||
- **Threadripper PRO (3000WX/5000WX)** doubles the lane count to 128 but at meaningfully higher used
|
||||
cost (workstation-tier, lower volume, sWRX8 boards are pricier than sTRX4/TRX40 boards) — worth it
|
||||
only if 6 full-bandwidth (x16) GPU slots are actually needed; at x8-per-card (adequate for inference)
|
||||
64 lanes already covers 6 GPUs with lanes to spare for NVMe/chipset.
|
||||
|
||||
**Lane budget for 6 GPUs on sTRX4 (64 lanes), estimated (no vendor spec gives a topology this
|
||||
specific — treat this bullet as an estimate):** typical sTRX4 boards reserve ~4 lanes for the
|
||||
chipset uplink and commonly wire 1-2 M.2 slots directly to the CPU (4 lanes each) — so realistic
|
||||
GPU-available lanes land around 44-52 of the 64, i.e. **6 GPUs at x8 electrical each (48 lanes) is
|
||||
plausible but board-model-dependent**; x16-each for 6 cards is not possible on 64 lanes regardless of
|
||||
board. x8 electrical is not a meaningful inference-speed penalty for llama.cpp (weights are loaded
|
||||
once; the ongoing per-token traffic across PCIe is small compared to compute), so this is an
|
||||
acceptable tradeoff, not a real bottleneck for this workload.
|
||||
|
||||
---
|
||||
|
||||
## 6. Power budget vs. the ~200W / one spare 8-pin headroom
|
||||
|
||||
Official/vendor TDPs:
|
||||
|
||||
| Card | TDP | Source |
|
||||
|---|---|---|
|
||||
| GTX 1080 (owned) | 180W, one 8-pin | [confirmed 180W, PCIe 3.0 x16, 1× 8-pin](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu) — spec matches NVIDIA's own launch figures reported across multiple outlets incl. Tom's Hardware |
|
||||
| RTX 2080 (owned) | 215W | Cross-checked across gpuzoo/cputronic/notebookcheck spec pages, consistent at 215W |
|
||||
| GT 710 (owned) | ~19W, **no external power connector** (slot power only) | [MSI/EVGA/Zotac GT 710 spec pages](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification) |
|
||||
| RTX 3060 12GB (candidate purchase) | 170W, one 8-pin | [NVIDIA-confirmed 170W TDP, one 8-pin connector](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/) |
|
||||
| RTX 3090 24GB (candidate purchase) | 350W, two 8-pin, [NVIDIA's own RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/) lists 350W and a 750W PSU minimum | NVIDIA official |
|
||||
| R9700 32GB (already deployed / "more of the same") | 300W (per this repo's `server-planing.md`, consistent with AMD's own R9700 product page framing it as a 300W-class card) | in-repo prior research |
|
||||
|
||||
**Against the stated ~200W / one spare 8-pin budget:**
|
||||
|
||||
- Adding **one RTX 3060 12GB** (170W, one 8-pin) is the *only* candidate in this list that fits the
|
||||
stated headroom as-is — it uses the one spare connector and stays under 200W.
|
||||
- Adding the already-owned **GTX 1080** (180W) as the classifier's dedicated card also just barely
|
||||
fits (180W ≤ 200W, one 8-pin) — this is a genuinely free option since the card is already owned and
|
||||
its power draw is within budget, unlike every purchase candidate below.
|
||||
- Adding the already-owned **RTX 2080** (215W) **exceeds** the stated 200W headroom by 15W — technically
|
||||
over budget on paper, though real-world draw is usually a bit under rated TDP; flag it as marginal,
|
||||
not safely fitting.
|
||||
- Adding a **second R9700** (300W) or an **RTX 3090** (350W, needs two 8-pin — the user has only one
|
||||
spare) both blow well past the current power budget on both watts and connector count.
|
||||
- **The GT 710 draws no meaningful power (~19W, no PCIe power connector at all)** — it is free from a
|
||||
power-budget standpoint regardless of what else is added.
|
||||
|
||||
**PSU upgrade trigger:** the very first stage that adds *any* GPU beyond a GTX 1080-class card (180W,
|
||||
one 8-pin) or an RTX 3060 12GB (170W, one 8-pin) exhausts the stated headroom. Any stage that reaches
|
||||
for a second 32GB-class card (R9700 or equivalent) or any 300W+ card **requires a PSU upgrade before
|
||||
that stage**, not after — see the roadmap table in §7 for exactly which stage that is.
|
||||
|
||||
---
|
||||
|
||||
## 7. Mixed-GPU feasibility: ROCm + CUDA, and is the GT 710 usable at all
|
||||
|
||||
**ROCm and CUDA are different llama.cpp builds, but that's exactly the pattern already in this
|
||||
repo.** `ghcr.io/ggml-org/llama.cpp` publishes both `server-rocm` and `server-cuda` as separate,
|
||||
independently-built image tags (confirmed present on the [ggml-org container registry](https://github.com/orgs/ggml-org/packages/container/llama.cpp)
|
||||
and documented in [`docs/docker.md`](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md)
|
||||
— *"server-cuda: Same as `server` but compiled with CUDA support"*, *"server-rocm: Same as `server`
|
||||
but compiled with ROCm support"*). You cannot mix backends inside one process/container, but you
|
||||
**can** run one `server-rocm` container pinned to the R9700 and a separate `server-cuda` container
|
||||
pinned to an NVIDIA card, simultaneously, on the same host — this is architecturally identical to
|
||||
today's `llama-server` + `qwen-classifier` two-container split in `docker-compose.yml`, just with a
|
||||
different image tag for the NVIDIA-backed service and NVIDIA's container runtime (`nvidia-container-toolkit`
|
||||
+ `--gpus` / device reservation, the CUDA-world equivalent of this repo's `/dev/kfd`+`/dev/dri`+
|
||||
numeric-GID ROCm pattern documented in
|
||||
[`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md)). None of that doc's ROCm-specific
|
||||
findings (the `GPU_MAX_HW_QUEUES=1` MES firmware workaround, the numeric-GID `group_add` fix) apply to
|
||||
an NVIDIA/CUDA container — those are ROCm-stack-specific bugs, not general multi-GPU-container issues.
|
||||
|
||||
**Is this an implicit AMD→NVIDIA rebuild, or additive?** Worth surfacing explicitly since the two
|
||||
source plans conflict on this: `server-planing.md` is an AMD-only, ROCm-only 4×R9700 plan. This
|
||||
ticket's plan is **additive/mixed** — keep the R9700 running the main model under ROCm, and bolt on
|
||||
NVIDIA cards under CUDA for secondary roles (classifier, or a second inference GPU for the big model
|
||||
if going the "more of the same type" route means buying NVIDIA instead of more R9700s). Both are
|
||||
internally consistent, but they are different end-states — flag this choice back to the user rather
|
||||
than assuming one.
|
||||
|
||||
**Splitting the *main* 27B model itself across mixed AMD+NVIDIA silicon in one process is not
|
||||
possible** — llama.cpp's multi-GPU tensor-split only works within a single backend build. To use
|
||||
both an R9700 and an NVIDIA card for the *same* model's layers, all the compute-hosting cards need to
|
||||
be the same backend (all-ROCm or all-CUDA) in that one process. This is why §5's roadmap treats "add
|
||||
GPU capacity to the main model" and "add a GPU for the classifier" as separable purchases with
|
||||
different backend constraints, not a single mixed pool.
|
||||
|
||||
**Is the GT 710 usable for anything in this pipeline? No.** Reasoning:
|
||||
|
||||
- 1GB VRAM cannot hold any meaningful fraction of either model's weights (17.6 GB / 2.4-4.3 GB) —
|
||||
even a handful of transformer layers at Q4 quantization exceeds 1GB.
|
||||
- It's Kepler-generation silicon (192 CUDA cores, no tensor cores) — llama.cpp's CUDA backend
|
||||
technically supports pre-Turing cards, but at this VRAM size there's nothing to usefully offload.
|
||||
- It draws power from the PCIe slot only, no external connector — genuinely free to keep installed.
|
||||
- **Plausible actual use: dedicate it as the box's display-output card**, so every compute-capable
|
||||
GPU (R9700, and whichever NVIDIA cards get added) can be fully headless/compute-only with none of
|
||||
their VRAM or a display output tied up driving a monitor — a real, if minor, use for it. This is
|
||||
this doc's own inference from the spec facts above, not a claim found in any primary source.
|
||||
|
||||
---
|
||||
|
||||
## 8. GPU market pricing vs. the $20-30/GB assumption
|
||||
|
||||
| Card | VRAM | Backend | Current used-market price (estimate — see caveat) | $/GB |
|
||||
|---|---|---|---|---|
|
||||
| RTX 3060 12GB | 12GB | CUDA | ~$240-300 used (eBay listings, [gpupoet.com tracker](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060): *"from $239"*, [eBay live listings](https://www.ebay.com/shop/rtx-3060-12gb) averaging ~$488 asking but with a $239 floor) | **~$20-25/GB** — matches the stated assumption |
|
||||
| RTX 3090 24GB | 24GB | CUDA | ~$1,010-1,050 used ([bestvaluegpu.com Sep 2026 tracker](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/), [xda-developers coverage](https://www.xda-developers.com/used-rtx-3090-still-best-for-local-ai-in-value/)) | **~$42-44/GB** — well above the stated assumption |
|
||||
| R9700 32GB ("more of the same type") | 32GB | ROCm | **New only — $1,299 MSRP**, street price $1,400-1,585 as of this research ([overclock3d](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/), [pricehistory.app tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm)) — too recent a release (2026) for a used market to exist yet | **~$41-50/GB, and not a used-market price at all** |
|
||||
|
||||
**Caveat on all three price figures:** these are live marketplace asking-price snapshots pulled via
|
||||
web search on 2026-09-11, not sold-price data or a vendor spec sheet — treat as directional, not
|
||||
exact. eBay asking prices in particular run above realized sale prices.
|
||||
|
||||
**Correction to the user's stated assumption:** $20-30/GB is a good estimate specifically for
|
||||
**last-generation mainstream used cards** (RTX 3060 12GB fits it almost exactly) but **not** for
|
||||
high-VRAM flagship cards like the RTX 3090 (~1.5-2× that rate) or for "more of the same type" R9700
|
||||
units, which aren't used-market at all yet and sit even higher per GB than the 3090. If the plan is
|
||||
"cheapest path to more VRAM," multiple RTX 3060 12GB cards (or similar mid-tier used cards) beat one
|
||||
RTX 3090 on $/GB, at the cost of needing more PCIe slots and more total wattage/connectors to reach
|
||||
the same aggregate VRAM — which is exactly the tradeoff the Threadripper lane-count plan in §5 is
|
||||
for.
|
||||
|
||||
---
|
||||
|
||||
## 9. Step-by-step roadmap
|
||||
|
||||
All "resulting max context" figures assume `--parallel 2` and the KV precision stated; per §2, the
|
||||
`--ctx-size` value shown is the *total* (pre-division) value to pass to llama-server.
|
||||
|
||||
| Stage | Hardware change | Est. cost | Backend | Usable VRAM (main-model pool) | Max context @ parallel=2 (`q4_0` KV) | PSU upgrade triggered? |
|
||||
|---|---|---|---|---|---|---|
|
||||
| **0 (current)** | 1× R9700 32GB, in production | $0 | ROCm | 32GB (shared with classifier) | ~131K/slot today at `q8_0` KV (262144 total ÷ 2) | No |
|
||||
| **1 — classifier isolation** | Move classifier onto the already-owned **GTX 1080** (180W, own container, `server-cuda`), freeing the R9700 entirely for the main model. Matches the existing dual-model pattern qwen-code's own docs describe (§10) and this repo's `qwen-classifier` service already implements, just on separate silicon instead of a shared card. | $0 (already owned) | ROCm (main) + CUDA (classifier) | R9700's full 32GB now available to the main model alone | ~262K/slot @ `q8_0` (unchanged ctx-size, no more classifier contention) | **No** — 180W GTX 1080 fits the stated ~200W/one-8-pin headroom |
|
||||
| **2 — second big-model GPU** | Add **one more 32GB-class card** for the main model. Cheapest correct-backend option: a second R9700 (~$1,300-1,585 new, ROCm, same backend as the first — required if tensor-splitting one model across two cards) | ~$1,300-1,585 | ROCm | 64GB pooled | `--ctx-size 500000 --parallel 1` fits at `q4_0` (~33GB) or `q8_0` (~48GB, tight but fits in 64GB) — **not yet 500k×2** | **Yes** — 300W card, no spare 8-pin left after stage 1 |
|
||||
| **3 — reach 500k × 2 / 1M stretch** | No further hardware if stage 2's 64GB pool is used with `--cache-type-k/v q4_0`: `--ctx-size 1000000 --parallel 2` needs ~33-39GB (§4), fits inside 64GB with real headroom for the compute buffer. If `q8_0` KV is required instead (this repo's current quality bar for the main model), the ~51-54GB need is tight-to-marginal on 64GB — a **third** 32GB card (~96GB pool) removes the risk. | $0 (reuses stage 2) or +$1,300-1,585 for a 3rd card if `q8_0` KV is required | ROCm | 64GB (q4_0 case) or 96GB (q8_0 case) | **500k×2 parallel achieved**, and the 1M stretch goal is the *same config* with `--parallel 1` instead of 2 (§2) | Already upgraded at stage 2 |
|
||||
| **4 — optional CPU/lane platform swap** | Only needed if the plan is to keep scaling past 2-3 big cards, or to add several small used cards (RTX 3060 12GB) for extra headroom/throughput rather than raw ctx-size. Swap to **non-PRO Threadripper 3000-series (sTRX4)** — 64 PCIe 4.0 lanes, ~x8-per-slot for up to 6 GPUs (§5). Threadripper PRO 3000WX/5000WX (128 lanes) only if x16-per-card matters or 6+ full-bandwidth slots are wanted. | Used sTRX4 CPU+board: roughly $400-800 combined on the used market (not independently priced in this pass — **estimate**, not cited) | n/a (platform only) | n/a | n/a | Independent of GPU wattage — driven by whatever GPU count/wattage stage 5+ adds |
|
||||
| **5+ — scale-out via small used cards** | Add RTX 3060 12GB units (~$20-25/GB, the assumption that actually holds, §8) instead of more 32GB flagship cards, once lane count (stage 4) supports it — useful for extra parallel slots / throughput beyond the 500k×2 target rather than for raising ctx-size further (500k×2/1M is already met at stage 3). | ~$240-300/card | CUDA (separate container per §6) | +12GB pooled per card, but on a *different backend* from the ROCm main model — usable for extra classifier/small-model capacity or a separate CUDA-backend llama-server instance, not as additional tensor-split VRAM for the ROCm main model | Unchanged for the main model; adds parallel capacity elsewhere | Yes, cumulative — each additional 170W card needs PSU headroom stage 2 already consumed |
|
||||
|
||||
**Where the existing dual-model pattern sits in this roadmap:** it's stage 1, and it's free. The
|
||||
qwen-code docs pattern (main model + a small, always-resident, non-thinking fast/classifier model —
|
||||
see §10) is already implemented in this repo; the only roadmap-relevant change is *which GPU* the
|
||||
classifier sits on, moving it off the R9700 entirely onto an already-owned NVIDIA card frees the
|
||||
R9700's full 32GB for the 500k×2/1M push instead of splitting it with the classifier as happens
|
||||
today.
|
||||
|
||||
### 9.1 Upgrade path, as diagrams
|
||||
|
||||
Diagram form of the same §9 table and §11.9's dense-vs-Flash-Next call — nothing new is claimed here,
|
||||
this is a visual index back into the cited sections above.
|
||||
|
||||
**Stage-by-stage hardware path** (PSU-upgrade triggers and target reached called out inline):
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
S0["Stage 0 — today<br/>1x R9700 32GB, ROCm<br/>classifier shares the card<br/>$0"]
|
||||
S1["Stage 1 — classifier isolation<br/>+ GTX 1080 (owned, 180W, CUDA)<br/>R9700 freed for main model<br/>$0 · PSU OK (180W fits ~200W headroom)"]
|
||||
S2["Stage 2 — 2nd big-model GPU<br/>+1x R9700 32GB (ROCm)<br/>64GB pooled<br/>~$1,300-1,585 · PSU UPGRADE REQUIRED (300W, no 8-pin left)"]
|
||||
S3q4["Stage 3a — q4_0 KV<br/>--ctx-size 1,000,000 --parallel 2<br/>~33-39GB, fits in 64GB<br/>$0 (reuses stage 2)"]
|
||||
S3q8["Stage 3b — q8_0 KV (current prod quality)<br/>~51-54GB, tight on 64GB<br/>+1x R9700 -> 96GB removes risk<br/>+~$1,300-1,585"]
|
||||
TARGET(["500k x2 parallel reached<br/>= 1M stretch goal, same VRAM<br/>(--parallel 1 vs 2 is a config flag, §2)"])
|
||||
S4["Stage 4 — platform swap (optional)<br/>sTRX4 Threadripper 3000, 64 PCIe4 lanes<br/>only needed past 2-3 big cards<br/>~$400-800 (estimate, §9)"]
|
||||
S5["Stage 5+ — scale out<br/>+RTX 3060 12GB cards (CUDA, separate backend)<br/>extra parallel/throughput, not more ctx-size<br/>~$240-300/card · PSU upgrade each card"]
|
||||
|
||||
S0 --> S1 --> S2
|
||||
S2 --> S3q4 --> TARGET
|
||||
S2 --> S3q8 --> TARGET
|
||||
TARGET -.->|"only if scaling past this"| S4 --> S5
|
||||
|
||||
style TARGET fill:#2e7d32,color:#fff,stroke:#1b5e20
|
||||
style S2 fill:#8a5a00,color:#fff,stroke:#5c3d00
|
||||
style S5 fill:#8a5a00,color:#fff,stroke:#5c3d00
|
||||
```
|
||||
|
||||
**Model choice, and the one open question that could change it** (§11.9):
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Q{"Goal: 500k-1M ctx<br/>within a $20-30/GB VRAM budget?"}
|
||||
D["Dense Qwen3.8-27B<br/>17.6-54.7GB weights (Q4_K_XL-BF16)<br/>reaches 500k x2 on 2-3 cards,<br/>500k x4 on 2-6 cards depending on quant<br/>(§11.7 tables)"]
|
||||
F{"Try --n-cpu-moe:<br/>offload MoE experts to system RAM?<br/>(untested for this model, §11.8)"}
|
||||
FBAD["Flash-Next, all-GPU weights<br/>111-354GB just for weights<br/>needs 4-13 cards before any KV cost<br/>NOT recommended at this budget (§11.9)"]
|
||||
FGOOD["Flash-Next, experts in system RAM<br/>GPU VRAM could shrink a lot<br/>2.67x cheaper KV/token becomes relevant<br/>UNVERIFIED — prototype on real server first"]
|
||||
CAVEAT["+ real caveat either way:<br/>PR #27742 flags unverified conv branch,<br/>3% QSA divergence, prefill-pos-0-only PLE<br/>(§11.1) — dense model carries no such flag"]
|
||||
|
||||
Q --> D
|
||||
Q -->|"considering Flash-Next instead"| F
|
||||
F -->|"works well"| FGOOD
|
||||
F -->|"doesn't help / untested"| FBAD
|
||||
FGOOD --> CAVEAT
|
||||
FBAD --> CAVEAT
|
||||
|
||||
style D fill:#2e7d32,color:#fff,stroke:#1b5e20
|
||||
style FBAD fill:#8a1c1c,color:#fff,stroke:#5c1212
|
||||
style FGOOD fill:#8a5a00,color:#fff,stroke:#5c3d00
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 10. Qwen-code's own docs on the fast-model/classifier pattern
|
||||
|
||||
Fetched directly per the user's link:
|
||||
[qwenlm.github.io/qwen-code-docs/en/users/overview/](https://qwenlm.github.io/qwen-code-docs/en/users/overview/)
|
||||
— **the overview page itself does not describe the dual-model/classifier pattern**; it only covers
|
||||
single-model-provider setup (Alibaba ModelStudio / third-party / custom provider), one model at a
|
||||
time. The actual fast-model/classifier documentation lives on the **Auto Mode** page instead, which
|
||||
this repo's own `fast-model-choice.md` already fetched and cited in detail:
|
||||
[qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/) —
|
||||
summary (see `fast-model-choice.md` §1 for the full quote): a two-stage classifier gate, both stages
|
||||
using "your configured fast model (`/model --fast`)", Stage 1 a ~300ms `{shouldBlock}`-only check,
|
||||
Stage 2 a ~3-5s chain-of-thought review that only runs on a Stage-1 block. Nothing in either page
|
||||
gives a recommended *context size* or *model size* for the fast model beyond what's implied by that
|
||||
latency budget — this repo's own prior research (`fast-model-choice.md`) derived the actual context
|
||||
requirement from qwen-code's source code instead (`packages/core/src/permissions/classifier-transcript.ts`),
|
||||
since the docs pages don't state one. No new information changes that prior doc's conclusion; this
|
||||
section exists to confirm the overview page was checked directly as instructed and doesn't contradict
|
||||
or add to it.
|
||||
|
||||
---
|
||||
|
||||
## 11. Alternative: Qwen3.8-Flash-Next (MoE, hybrid attention)
|
||||
|
||||
The user also wants to weigh switching (or adding) **Qwen3.8-Flash-Next** — a 125B-total/6B-active MoE
|
||||
with a hybrid recurrent-attention architecture — against staying on dense Qwen3.8-27B, and separately
|
||||
wants this section to cover **going up in weight quant** (Q4_K_XL → Q6_K_XL → Q8_0 → BF16/fp16) for
|
||||
*both* models, not just Q4. Feasibility first, since it gates everything else.
|
||||
|
||||
### 11.1 Feasibility verdict: supported, but immature — read before trusting any number below
|
||||
|
||||
Checked directly against the primary sources the task named:
|
||||
|
||||
- **llama.cpp mainline support exists.** [PR #27742](https://github.com/ggml-org/llama.cpp/pull/27742)
|
||||
("model: add Qwen3.8-Flash-Next (qwen4exp)") was **merged into `master` on 2026-08-27** by ngxson.
|
||||
It adds the full architecture: Gated DeltaNet layers (sigmoid-gated linear attention), QSA
|
||||
("Qwen Sparse Attention", operating at micro-block granularity), hyper-connections, and the PLE
|
||||
n-gram embedding table. `llama.cpp`'s own docs list CPU/CUDA/Metal/ROCm as supported backends for
|
||||
it — this is not a CUDA-only feature.
|
||||
- **This repo's pinned image is a floating tag, not a version pin.** `docker-compose.yml` runs
|
||||
`ghcr.io/ggml-org/llama.cpp:server-rocm` with no date/digest suffix — a rolling "latest ROCm server
|
||||
build" tag, not a release version. The merge is from 2026-08-27, and today is 2026-09-11 (~2 weeks
|
||||
later), so a **fresh pull** of `server-rocm` should include it — but whatever image is already
|
||||
cached/running on the R9700 box may predate the merge. **Action before touching this model on the
|
||||
server: `docker compose pull llama-server` and check the startup log's build/commit banner is dated
|
||||
on/after 2026-08-27**, not just "the tag says server-rocm."
|
||||
- **Real, primary-source-flagged immaturity — this is the part that should temper enthusiasm.** The
|
||||
PR's own description/review discussion states: *"The conv branch itself is still numerically
|
||||
unverified because the fixture zeroes its weights"*; QSA sparse attention *"diverges on 3 percent of
|
||||
positions"* above its budget threshold; the PLE depthwise convolution *"is exact only for a prefill
|
||||
that starts at position 0"* (i.e. correctness is not guaranteed once `--cache-reuse`/prompt-caching
|
||||
is in play — a flag this repo already turns on for the dense model per the latest commit). None of
|
||||
that is disqualifying, but it is a primary-source admission that this is a fresh, not-fully-verified
|
||||
implementation, not a mature, widely-battle-tested one like the dense Qwen3.8-27B path.
|
||||
- **Multi-slot serving needs an explicit new flag.** The same PR states: *"`set_input_qsa` asserted
|
||||
`n_stream == 1`, so llama-server could not serve this model with more than one slot unless `-kvu`
|
||||
was passed."* Per the server README (§2), `--kv-unified`/`-kvu` defaults to enabled **only when slot
|
||||
count is auto** (`-1`). This repo's compose file sets `--parallel ${LLAMA_PARALLEL:-2}` **explicitly**
|
||||
(not auto) — so adopting Flash-Next with `--parallel` > 1 requires **adding `--kv-unified` (or
|
||||
`-kvu`) to the launch flags**, a real deploy-time change, not something that "just works" by copying
|
||||
today's flag set onto a new model file.
|
||||
|
||||
**Verdict: yes, runnable** on this repo's backend (ROCm, mainline, no dev branch needed) as long as the
|
||||
image is pulled after 2026-08-27 and `-kvu` is added for multi-slot use — but treat it as
|
||||
**usable-with-caution**, not a drop-in swap, given the PR author's own unresolved-correctness notes.
|
||||
|
||||
### 11.2 Architecture, verified against `config.json` directly
|
||||
|
||||
Fetched from `Qwen/Qwen3.8-Flash-Next`'s `config.json` (unsloth's GGUF repo repackages the same base
|
||||
model): `num_hidden_layers=48`, `hidden_size=2560`, `num_attention_heads=24`, `num_key_value_heads=2`,
|
||||
`head_dim=256`, `max_position_embeddings=262144` (same native/extensible-to-1M framing as the dense
|
||||
model — same YaRN quality caveat from §3 applies here too, unverified past native length), `num_experts=512`,
|
||||
`num_experts_per_tok=10`, and the linear-attention head config: `linear_num_key_heads=16`,
|
||||
`linear_num_value_heads=48`, `linear_key_head_dim=128`, `linear_value_head_dim=128`.
|
||||
|
||||
Layer pattern (confirmed both from the model card's own description and `config.json`'s
|
||||
`full_attention_interval=4`): every 4th layer is full/QSA attention, the other 3 are Gated DeltaNet —
|
||||
**12 of 48 layers grow a real KV cache; the other 36 have a fixed-size recurrent state that does not
|
||||
grow with context length.** (24% full-attention layers vs. the dense model's 16-of-64 = 25% — similar
|
||||
ratio, but the *absolute* per-layer KV cost differs because `num_key_value_heads` is 2 here vs. 4 on
|
||||
the dense model — see below.)
|
||||
|
||||
### 11.3 Per-token growing-KV-cache cost
|
||||
|
||||
```
|
||||
12 full-attention layers × 2 (K+V) × 2 kv_heads × 256 head_dim × 2 bytes (fp16) = 24 KiB/token
|
||||
```
|
||||
|
||||
| Total ctx-size | KV cache, fp16 | KV cache, q8_0 | KV cache, q4_0 |
|
||||
|---|---|---|---|
|
||||
| 262,144 (native) | ~6.0 GiB | ~3.0 GiB | ~1.5 GiB |
|
||||
| 500,000 | ~11.4 GiB | ~5.7 GiB | ~2.9 GiB |
|
||||
| **1,000,000 (500k×2, or 1M stretch)** | **~22.9 GiB** | **~11.4 GiB** | **~5.7 GiB** |
|
||||
| **2,000,000 (500k×4)** | **~45.8 GiB** | **~22.9 GiB** | **~11.4 GiB** |
|
||||
|
||||
`--cache-type-k/v` are the same generic llama.cpp KV-cache-quantization flags used elsewhere in this
|
||||
doc; nothing in the PR or the server README suggests they're handled differently for the 12
|
||||
full-attention layers of a hybrid model — they quantize the same growing K/V buffers as on a plain
|
||||
transformer. (No primary source explicitly confirms this for *this* architecture specifically — flagged
|
||||
as a reasonable extrapolation, not a directly-cited fact, same caveat class as this doc's other
|
||||
estimates.)
|
||||
|
||||
### 11.4 Fixed (non-growing) recurrent state — Gated DeltaNet layers
|
||||
|
||||
The 36 Gated DeltaNet layers each keep a fixed-size recurrent state (an outer-product-style
|
||||
key×value matrix per head) that does **not** scale with context length — only with slot/sequence
|
||||
count. Sized from `config.json`'s linear-attention head params:
|
||||
|
||||
```
|
||||
36 layers × linear_num_value_heads(48) × linear_key_head_dim(128) × linear_value_head_dim(128) × 4 bytes (fp32 state)
|
||||
≈ 36 × 48 × 128 × 128 × 4 bytes ≈ 108 MiB per slot
|
||||
```
|
||||
|
||||
This is **this doc's own derivation from the published head-dimension params, not a value pulled
|
||||
directly from llama.cpp source or docs** — the PR text confirms the state exists per-stream/per-slot
|
||||
but doesn't publish an exact byte formula, so treat the ~108 MiB/slot figure as an estimate, medium
|
||||
confidence. Even at 4 parallel slots that's under half a gigabyte — **negligible** next to both the
|
||||
growing KV cache (GBs) and the weights (tens to hundreds of GB) computed below. The headline
|
||||
implication holds regardless of the exact multiplier: Flash-Next's "big memory line item" is the MoE
|
||||
weights, not the attention state of any kind.
|
||||
|
||||
### 11.5 Magnitude vs. the dense model — how much cheaper is KV, really
|
||||
|
||||
At the same total ctx-size, Flash-Next's growing KV cache is **24 KiB/token vs. the dense model's
|
||||
64 KiB/token — 2.67× smaller**, i.e. Flash-Next's KV budget is **37.5%** of the dense model's at
|
||||
identical context length. This is a real, significant win *for the KV-cache line item specifically* —
|
||||
but see §11.9: it's a much smaller slice of a much bigger total, because the weights move the other
|
||||
way by a far larger factor.
|
||||
|
||||
### 11.6 Weight sizes — verified from each unsloth GGUF repo's actual file listing
|
||||
|
||||
Fetched directly from the HF file trees (not estimated from ratios), current as of this research pass:
|
||||
|
||||
| Quant tier | Qwen3.8-27B (dense) | Qwen3.8-Flash-Next (MoE) |
|
||||
|---|---|---|
|
||||
| Q4_K_XL (`UD-Q4_K_XL`) | **17.6 GB** (existing baseline) | **111.4 GB** (4 parts: 10.9MB + 49.9GB + 49.4GB + 12.1GB) |
|
||||
| Q6_K_XL (`UD-Q6_K_XL`) | **25.3 GB** | **169 GB** (6 parts) |
|
||||
| Q8_0 | **29 GB** | **188 GB** (6 parts) |
|
||||
| BF16/fp16 | **54.67 GB** (50GB + 4.67GB, 2 parts) | **354 GB** (8 parts) |
|
||||
|
||||
Sources: [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main)
|
||||
and its `BF16/` subfolder; [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main)
|
||||
and its `UD-Q4_K_XL/`, `UD-Q6_K_XL/`, `Q8_0/`, `BF16/` subfolders (per-file sizes summed). The
|
||||
preliminary Q8_0 figure floated before this research pass (~192GB) was slightly high — the real
|
||||
listing sums to **188 GB**; everything else in the preliminary list was accurate to within rounding.
|
||||
|
||||
**The weight-quant axis and the KV-cache-quant axis are independent knobs.** Raising weight quality
|
||||
(Q4_K_XL → BF16) does not require raising `--cache-type-k/v` — the two flags are unrelated, and this
|
||||
repo already proves that pattern works (`q8_0` KV cache is deployed today against `Q4_K_XL` weights).
|
||||
A user chasing **maximum output quality** can run e.g. **BF16 weights + `q4_0` KV cache** — full-precision
|
||||
weights for quality, still-compressed KV for context budget — or any other combination in the tables
|
||||
below; nothing about picking a higher weight quant forces a matching KV precision.
|
||||
|
||||
### 11.7 Total VRAM: does it fit, across quant tiers and both parallelism targets
|
||||
|
||||
All totals = weights + growing KV cache + an estimated **+3-6 GB** compute-buffer/runtime overhead
|
||||
(same estimate band as §4, carried over — not re-derived for this architecture; flagged medium
|
||||
confidence there too). "Cards" = ceil(total ÷ 32GB), i.e. how many R9700-class 32GB cards it takes.
|
||||
|
||||
#### Dense Qwen3.8-27B — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`)
|
||||
|
||||
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|
||||
|---|---|---|---|
|
||||
| Q4_K_XL (17.6GB) | ~79-85GB (**3**) | ~48-54GB (**2**) | ~33-39GB (**2**) |
|
||||
| Q6_K_XL (25.3GB) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) | ~44-47GB (**2**) |
|
||||
| Q8_0 (29GB) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) | ~47-50GB (**2**) |
|
||||
| BF16 (54.67GB) | ~119-122GB (**4**) | ~88-91GB (**3**) | ~73-76GB (**3**) |
|
||||
|
||||
#### Dense Qwen3.8-27B — 500k × 4 parallel (`--ctx-size 2,000,000`)
|
||||
|
||||
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|
||||
|---|---|---|---|
|
||||
| Q4_K_XL (17.6GB) | ~143-146GB (**5**) | ~82-85GB (**3**) | ~51-54GB (**2**) |
|
||||
| Q6_K_XL (25.3GB) | ~150-153GB (**5**) | ~89-92GB (**3**) | ~59-62GB (**2**, tight) |
|
||||
| Q8_0 (29GB) | ~154-157GB (**5**) | ~93-96GB (**3**, edge) | ~63-66GB (**2-3**, edge) |
|
||||
| BF16 (54.67GB) | ~180-183GB (**6**) | ~119-122GB (**4**) | ~88-91GB (**3**) |
|
||||
|
||||
#### Qwen3.8-Flash-Next — 500k × 2 parallel / 1M stretch (`--ctx-size 1,000,000`)
|
||||
|
||||
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|
||||
|---|---|---|---|
|
||||
| UD-Q4_K_XL (111.4GB) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) | ~120-123GB (**4**) |
|
||||
| UD-Q6_K_XL (169GB) | ~195-198GB (**7**) | ~183-186GB (**6**) | ~178-181GB (**6**) |
|
||||
| Q8_0 (188GB) | ~214-217GB (**7**) | ~202-205GB (**7**) | ~197-200GB (**7**) |
|
||||
| BF16 (354GB) | ~357-360GB (**12**) | ~357-360GB (**12**) | ~357-360GB (**12**) |
|
||||
|
||||
#### Qwen3.8-Flash-Next — 500k × 4 parallel (`--ctx-size 2,000,000`)
|
||||
|
||||
| Weight quant | fp16 KV → total (cards) | q8_0 KV → total (cards) | q4_0 KV → total (cards) |
|
||||
|---|---|---|---|
|
||||
| UD-Q4_K_XL (111.4GB) | ~160-163GB (**6**) | ~137-140GB (**5**) | ~126-129GB (**4**, edge) |
|
||||
| UD-Q6_K_XL (169GB) | ~218-221GB (**7**) | ~195-198GB (**7**) | ~183-186GB (**6**) |
|
||||
| Q8_0 (188GB) | ~234-237GB (**8**) | ~211-214GB (**7**) | ~199-202GB (**7**) |
|
||||
| BF16 (354GB) | ~397-400GB (**13**) | ~377-380GB (**12**) | ~366-369GB (**12**) |
|
||||
|
||||
(Flash-Next's KV precision barely moves the total at any weight quant above `UD-Q6_K_XL` — the weights
|
||||
so dominate the budget that KV quantization stops mattering for the "how many cards" question. This
|
||||
is the clearest signal in this whole section: for Flash-Next, the weight-quant choice is the entire
|
||||
hardware-sizing decision; for the dense model, KV precision still matters a lot.)
|
||||
|
||||
### 11.8 CPU MoE-expert offload — the one lever that could change this calculus
|
||||
|
||||
Flash-Next is a 512-expert/10-active-per-token MoE, and llama.cpp has a purpose-built flag for exactly
|
||||
this shape of model, confirmed directly from the server README: **`--n-cpu-moe`** — *"keep the Mixture
|
||||
of Experts (MoE) weights of the first N layers in the CPU"* — plus the more general
|
||||
**`--override-tensor`** (*"override tensor buffer type"*, pattern-matched by tensor name) that the same
|
||||
flag is built on top of. Both are generic, architecture-agnostic llama.cpp mechanisms (they match on
|
||||
tensor name patterns, not model type), so there's no reason to expect them not to apply to Flash-Next's
|
||||
MoE tensors specifically — but this pass found **no primary source that has actually tested
|
||||
`--n-cpu-moe` against this specific qwen4exp architecture**, so treat "it works here" as plausible,
|
||||
not confirmed.
|
||||
|
||||
If it does work as expected, this changes the whole weight-VRAM picture in §11.7: the ~90-95% of
|
||||
Flash-Next's weight footprint that's MoE expert tensors could live in system RAM while attention
|
||||
projections, the shared/non-expert tensors, and the full KV cache stay on GPU — meaning a much smaller
|
||||
GPU-VRAM number than the "all weights on GPU" tables above, at the cost of PCIe/RAM-bandwidth-bound
|
||||
inference speed for whichever experts get selected per token (this repo has no benchmark of that
|
||||
tradeoff, and it's highly system-RAM-bandwidth-dependent, so no number is given here — flagged as an
|
||||
escape hatch worth prototyping directly on the server, not something this research values responsibly
|
||||
without a real test run).
|
||||
|
||||
### 11.9 Net recommendation: dense Qwen3.8-27B vs. Flash-Next, for this user's stated goal
|
||||
|
||||
**Net loss for this user's goal, as things stand — stay on dense Qwen3.8-27B.** Reasoning:
|
||||
|
||||
- The user's target (500k×2 or 500k×4, on a $20-30/GB-VRAM budget, GPUs in 32GB increments) is a
|
||||
**VRAM-budget-constrained** goal, and §11.7 shows Flash-Next's *weights alone* (111-354GB depending
|
||||
on quant) dwarf the entire dense-model total-VRAM figure from §4/§11.7 (33-183GB depending on quant)
|
||||
at every parallelism target. Flash-Next's much cheaper per-token KV cache (§11.5, real and verified)
|
||||
is a rounding error next to that weight-size gap — the "2.67× cheaper KV" win doesn't come close to
|
||||
offsetting a "6-20× larger weight footprint," so at $20-30/GB-VRAM the *dense* model reaches 500k×2
|
||||
or 500k×4 for a fraction of the card count and dollar cost that Flash-Next needs even at its lowest
|
||||
usable quant (`UD-Q4_K_XL`, 4-5 cards minimum) — before even factoring in §11.1's immaturity flags.
|
||||
- The one scenario that could flip this verdict is `--n-cpu-moe` actually working well for this
|
||||
architecture (§11.8) — if most of those 111-354GB of expert weights can sit in system RAM at
|
||||
acceptable throughput, Flash-Next's GPU-VRAM number could shrink dramatically and its real KV-cache
|
||||
advantage would start to matter. That is untested here and shouldn't be assumed; it's the one
|
||||
concrete next step worth trying on the actual server before ruling Flash-Next out permanently.
|
||||
- Independent of VRAM: §11.1's primary-source-flagged correctness caveats (unverified conv branch,
|
||||
3%-divergence QSA, prefill-position-0-only PLE exactness) are a real quality/stability risk on a
|
||||
production coding-agent stack that dense Qwen3.8-27B simply doesn't carry, since it's been running
|
||||
in this repo already.
|
||||
|
||||
---
|
||||
|
||||
## 12. 4-parallel × 500k scenario — all four combinations side by side
|
||||
|
||||
Per §2's already-established, cited rule (`n_ctx_per_seq = n_ctx / n_parallel`,
|
||||
[ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681)), the same division
|
||||
applies at 4 slots: **500k tokens on each of 4 parallel slots requires `--ctx-size 2,000,000`** —
|
||||
double the 2-parallel target's `--ctx-size 1,000,000`, for the same reason 500k×2 needed double
|
||||
262,144. This isn't a new mechanism, just the same formula at `--parallel 4`.
|
||||
|
||||
Full per-quant-tier tables for all four combinations are in §11.7 above (dense×2, dense×4, Flash-Next×2,
|
||||
Flash-Next×4 are each their own table there). Headline comparison at the KV precision already proven
|
||||
in production in this repo (`q8_0`) and each model's respective current/cheapest-usable weight quant:
|
||||
|
||||
| Scenario | `--ctx-size` | Weight quant | q8_0-KV total VRAM | Cards (32GB) |
|
||||
|---|---|---|---|---|
|
||||
| Dense × 2 (or 1M stretch) | 1,000,000 | Q4_K_XL (17.6GB, current) | ~48-54GB | **2** |
|
||||
| Dense × 4 | 2,000,000 | Q4_K_XL (17.6GB, current) | ~82-85GB | **3** |
|
||||
| Flash-Next × 2 (or 1M stretch) | 1,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~126-129GB | **4**, edge |
|
||||
| Flash-Next × 4 | 2,000,000 | UD-Q4_K_XL (111.4GB, cheapest usable) | ~137-140GB | **5** |
|
||||
|
||||
**4-parallel × 500k reachability against this repo's existing roadmap stages (§9):**
|
||||
|
||||
- **(a) Current 1×R9700 32GB:** none of the four combinations fit — not even dense×2 at any weight/KV
|
||||
quant (§4's own conclusion, unchanged).
|
||||
- **(b) The 2-3×R9700 roadmap already proposed in §9 (64-96GB):** covers **dense×2 fully** (stage 3, as
|
||||
already established) and **dense×4 at `q4_0` KV with Q4_K_XL or Q6_K_XL weights** (~51-62GB, fits in
|
||||
64-96GB) — but **not** dense×4 at higher weight quants (Q8_0/BF16 need 3-6 cards depending on KV
|
||||
precision, per §11.7's dense×4 table) and **not any Flash-Next scenario** (minimum is 4 cards/128GB
|
||||
even at the cheapest usable quant and tightest KV).
|
||||
- **(c) The full 4-6×R9700 stretch scenario** (`server-planing.md`'s original plan, 128-192GB pooled):
|
||||
covers **dense×4 at every weight quant up to BF16** (worst case ~91GB at BF16+q4_0, well inside
|
||||
128GB) and **Flash-Next×2 at `UD-Q4_K_XL`** (126-140GB, fits a 5-card/160GB build, tight on a 4-card/
|
||||
128GB one) — but **not** Flash-Next×4 at any weight quant above `UD-Q4_K_XL`, and not Flash-Next at
|
||||
`BF16` under any parallelism (needs 12-13 cards, an entirely different scale of build than anything
|
||||
in this doc's roadmap).
|
||||
|
||||
---
|
||||
|
||||
## Sources
|
||||
|
||||
- [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md) — `--ctx-size`, `--parallel`, `--cache-type-k/v`, `--kv-unified`, `--cache-reuse`, `--n-cpu-moe`, `--override-tensor` flag definitions
|
||||
- [ggml-org/llama.cpp#11681](https://github.com/ggml-org/llama.cpp/issues/11681) — real server log proving `n_ctx_per_seq = n_ctx / n_parallel`
|
||||
- [ggml-org/llama.cpp#27742](https://github.com/ggml-org/llama.cpp/pull/27742) — "model: add Qwen3.8-Flash-Next (qwen4exp)", merged 2026-08-27; architecture details, `n_stream == 1` / `-kvu` multi-slot requirement, and the conv-branch/QSA-divergence/PLE-prefill correctness caveats
|
||||
- [Qwen/Qwen3.8-27B config.json](https://huggingface.co/Qwen/Qwen3.8-27B/raw/main/config.json)
|
||||
- [Qwen/Qwen3.8-Flash-Next config.json](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/raw/main/config.json)
|
||||
- [Qwen/Qwen3-4B-Instruct-2507 config.json](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507/raw/main/config.json)
|
||||
- [unsloth/Qwen3.8-27B-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/tree/main) — 17.6GB (Q4_K_XL), 25.3GB (Q6_K_XL), 29GB (Q8_0), 54.67GB (BF16) weight sizes
|
||||
- [unsloth/Qwen3.8-Flash-Next-GGUF file tree](https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main) — 111.4GB (UD-Q4_K_XL), 169GB (UD-Q6_K_XL), 188GB (Q8_0), 354GB (BF16) weight sizes, summed from each quant's per-file listing
|
||||
- [ggml-org/llama.cpp docs/docker.md](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/docs/docker.md) — `server-cuda`/`server-rocm` separate image tags
|
||||
- [AMD chipsets product page](https://www.amd.com/en/products/processors/chipsets.html)
|
||||
- [Tom's Hardware — Threadripper 3960X/3970X, sTRX4/TRX40 launch](https://www.tomshardware.com/news/amd-unveils-threadripper-3960x-and-3970x-ryzen-9-3950x-details-and-athlon-3000g/2)
|
||||
- [PCWorld — Threadripper PRO launch, 128 PCIe lanes](https://www.pcworld.com/article/393181/amd-threadripper-pro-has-64-cores-128-pcie-lanes-and-8-channel-memory-support.html)
|
||||
- [NVIDIA — GeForce RTX 3090 product page](https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/)
|
||||
- [Lowyat.net — RTX 3060 official 170W TDP](https://www.lowyat.net/2021/232659/nvidia-geforce-rtx-3060-specifications-now-official-includes-3584-cuda-cores-and-170w-tdp/)
|
||||
- [BuildMyServer — GTX 1080 180W/PCIe3.0/1×8-pin spec listing](https://buildmyserver.com/products/zotac-nvidia-geforce-gtx-1080-8gb-gddr5-180w-pcie-3-0-x16-double-wide-gpu)
|
||||
- [MSI — GT 710 1GD5 LP spec page](https://www.msi.com/Graphics-Card/GT-710-1GD5-LP/Specification)
|
||||
- [overclock3d — AMD Radeon AI PRO R9700 $1,299 MSRP](https://overclock3d.net/news/gpu-displays/amd-unveils-its-1299-radeon-ai-pro-r9700-32gb-workstation-gpu/)
|
||||
- [pricehistory.app — R9700 street price tracker](https://pricehistory.app/p/powercolor-amd-radeon-ai-pro-r9700-32gb-BFcRhGIm)
|
||||
- [bestvaluegpu.com — RTX 3090 used price tracker, Sep 2026](https://bestvaluegpu.com/history/new-and-used-rtx-3090-price-history-and-specs/)
|
||||
- [gpupoet.com — RTX 3060 12GB used listings](https://gpupoet.com/gpu/shop/nvidia-geforce-rtx-3060)
|
||||
- [Qwen Code docs — overview](https://qwenlm.github.io/qwen-code-docs/en/users/overview/)
|
||||
- [Qwen Code docs — Auto Mode (fast-model pattern)](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/)
|
||||
- This repo: [`docker-compose.yml`](../../docker-compose.yml), [`.env.example`](../../.env.example), [`docs/server-planing.md`](../server-planing.md), [`qwen3.8-27b-quant.md`](qwen3.8-27b-quant.md), [`qwen3.8-27b-tool-calling.md`](qwen3.8-27b-tool-calling.md), [`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md), [`fast-model-choice.md`](fast-model-choice.md)
|
||||
|
||||
## Confidence/uncertainty summary
|
||||
|
||||
- **High confidence:** the KV-cache-per-token formulas for both dense models and Flash-Next (computed
|
||||
directly from each model's own `config.json`, same method this repo's prior research already used
|
||||
and cross-checked); the `n_ctx_per_seq = n_ctx / n_parallel` division behavior (directly evidenced by
|
||||
a real server log in a llama.cpp GitHub issue, and independently already documented in this repo's
|
||||
own `.env.example`) — and confirmed to apply identically at `--parallel 4` since the mechanism is
|
||||
parallel-count-agnostic; official TDP figures for GTX 1080, RTX 2080, RTX 3060, RTX 3090, GT 710
|
||||
(each cross-checked against 2+ independent spec listings or the vendor's own product page); the
|
||||
TR4-is-PCIe3/sTRX4-is-PCIe4 generational split (direct launch-coverage quote); the existence of
|
||||
separate `server-cuda`/`server-rocm` llama.cpp image tags; Qwen3.8-Flash-Next's `config.json`
|
||||
architecture params and PR #27742's merge date/status and its own stated correctness caveats and
|
||||
`-kvu` multi-slot requirement (all directly quoted from the primary source); the weight file sizes
|
||||
for both models at all four quant tiers (summed directly from each HF repo's real file listing, not
|
||||
estimated).
|
||||
- **Medium confidence:** the compute-buffer/runtime-overhead estimate in §4/§11.7 (+3-6 GiB) —
|
||||
extrapolated from one in-repo incident's before/after numbers, not a llama.cpp-documented formula,
|
||||
and carried over to Flash-Next without re-derivation for its different architecture; the Gated
|
||||
DeltaNet fixed recurrent-state size in §11.4 (~108 MiB/slot) — this doc's own derivation from the
|
||||
published head-dimension config, not a value found in llama.cpp source or docs; whether
|
||||
`--cache-type-k/v` quantization applies identically to Flash-Next's 12 full-attention layers as it
|
||||
does to a plain transformer (reasonable extrapolation, not directly confirmed for this architecture);
|
||||
whether `--n-cpu-moe`/`--override-tensor` actually work against Flash-Next's specific MoE tensor
|
||||
layout (architecture-agnostic mechanism, but untested against this model by any primary source found);
|
||||
real-world PCIe lane availability for 6 GPUs on a specific sTRX4 board (§5) — no single board's exact
|
||||
lane map was fetched, this is a reasonable-but-unverified estimate from typical sTRX4 board behavior.
|
||||
- **Low confidence / explicitly estimated, not cited fact:** all used-GPU marketplace pricing (§8) —
|
||||
live asking-price snapshots from a single search pass, not sold-price data; the used sTRX4
|
||||
CPU+motherboard combo price in the roadmap's stage 4 (§9) — not researched at all in this pass,
|
||||
flagged as a placeholder estimate; whether YaRN-scaled 500k/1M context actually holds output
|
||||
quality for either Qwen3.8-27B or Qwen3.8-Flash-Next — no primary source (Qwen's own docs included)
|
||||
publishes long-context quality benchmarks past the 262,144 native length for either model, so this is
|
||||
a known-unknown carried forward from each model card's "YaRN-extensible" claim, not a verified
|
||||
capability; whether the specific `ghcr.io/ggml-org/llama.cpp:server-rocm` image currently cached on
|
||||
this repo's server actually postdates PR #27742's 2026-08-27 merge — not checked against the live
|
||||
server in this pass, flagged as an action item in §11.1 rather than a confirmed fact.
|
||||
@@ -0,0 +1,249 @@
|
||||
# Does the qwen-classifier need to match the main model's context window, and would upgrading it to Qwen3-8B or Qwen3-Coder-30B-A3B-Instruct fit on the current R9700?
|
||||
|
||||
**Date:** 2026-09-15
|
||||
**Question raised:** should `qwen-classifier` be resized to `LLAMA_CTX_SIZE / LLAMA_PARALLEL` (131072, matching the
|
||||
main model's per-slot context), should the classifier model itself move up to Qwen3-8B or
|
||||
Qwen3-Coder-30B-A3B-Instruct, and does that need a second GPU?
|
||||
|
||||
**Answer: No, no, and not for this reason.** qwen-code's own docs state no context-size requirement for
|
||||
`fastModel` at all — the "match the main model" premise doesn't come from any primary source. Separately, and
|
||||
independently of context size: neither Qwen3-8B nor Qwen3-Coder-30B-A3B-Instruct fits in the ~6.1GiB of VRAM
|
||||
actually free on the card today, at *any* context length, once weights alone are counted — this is a raw-VRAM
|
||||
problem, not a context-window problem, exactly matching the "qwen8b needs to offload more to the RAM" intuition
|
||||
in the request. A second GPU would solve the VRAM problem (and incidentally remove this repo's own
|
||||
`GPU_MAX_HW_QUEUES=1` ROCm#5706 workaround from applying to this pair), but isn't deployed hardware today —
|
||||
it's a rack-build/acquisition question, not a config change.
|
||||
|
||||
## 1. Does qwen-code require the fast/classifier model to match the main model's context window?
|
||||
|
||||
No — checked against the user's own three linked pages, fetched directly:
|
||||
|
||||
- **`fastModel` settings docs**: "Model used for generating prompt suggestions and speculative execution,"
|
||||
configurable via `inherit` (main model), `fast`, a model ID, or `authType:model-id`; "Leave empty to use the
|
||||
main model." The docs recommend "a smaller/faster model (e.g., `qwen3-coder-flash`) reduces latency and
|
||||
cost" — **no context-window size or capacity requirement is stated anywhere on this page.**
|
||||
Source: [qwen-code docs — Configuration / Settings, `#fastmodel`](https://qwenlm.github.io/qwen-code-docs/en/users/configuration/settings/#fastmodel)
|
||||
|
||||
- **Approval Mode / Auto Mode classifier docs**: describes what the classifier evaluates (shell commands,
|
||||
network calls, out-of-workspace edits) and its allow/block behavior, but **does not name a specific model or
|
||||
state any context-window requirement** — the only operational note is that "when the classifier API is
|
||||
unreachable, the action is blocked rather than allowed."
|
||||
Source: [qwen-code docs — Approval Mode, `#4-auto-mode---classifier-driven-approval`](https://qwenlm.github.io/qwen-code-docs/en/users/features/approval-mode/#4-auto-mode---classifier-driven-approval)
|
||||
|
||||
- **Auto Mode "How it works"**: confirms the two-stage design (Stage 1: ~300ms, `{shouldBlock}` only; Stage 2:
|
||||
chain-of-thought reconsideration, only on a Stage-1 block) and what data reaches the classifier — user text,
|
||||
assistant tool-use calls, and tool-specific projections (truncated edit content, fetch URLs, shell command
|
||||
text). **Tool results are explicitly never sent to the classifier.** It "uses your configured fast model
|
||||
(`/model --fast`)," falling back to the main session model only if none is set. **No statement anywhere
|
||||
requires or implies the fast model's context window match the main model's.**
|
||||
Source: [qwen-code docs — Auto Mode, `#how-it-works`](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/#how-it-works)
|
||||
|
||||
This confirms and sharpens what this repo's own `fast-model-choice.md` already found by reading qwen-code's
|
||||
source directly (`packages/core/src/permissions/classifier-transcript.ts`: `MAX_TRANSCRIPT_MESSAGES=40`,
|
||||
`MAX_HISTORICAL_ACTION_CHARS=4000`/message, worst case ~40-50K tokens, live-tested at 15,116 prompt tokens) —
|
||||
that doc already called the original `131072` in `settings.json` "copied from the main model's entry, not a
|
||||
real qwen-code requirement." The three docs pages fetched here add nothing that contradicts that: **there is no
|
||||
primary-source basis for `LLAMA_CTX_SIZE / LLAMA_PARALLEL` symmetry between the two models.** The current
|
||||
`65536` classifier ctx already carries ~1.5x margin over the real worst case.
|
||||
|
||||
## 2. VRAM math for Qwen3-8B and Qwen3-Coder-30B-A3B-Instruct as classifier candidates
|
||||
|
||||
Same method this repo already uses (`qwen3.8-27b-quant.md`, `fast-model-choice.md` §5): per-token KV cache =
|
||||
`layers × 2(K+V) × kv_heads × head_dim × bytes`, read directly from each model's own `config.json`.
|
||||
|
||||
### Qwen3-8B
|
||||
|
||||
- Architecture (`Qwen/Qwen3-8B` `config.json`): `num_hidden_layers: 36`, `num_key_value_heads: 8`,
|
||||
`num_attention_heads: 32`, `head_dim: 128`, `hidden_size: 4096`, `max_position_embeddings: 40960`,
|
||||
`rope_scaling: null`.
|
||||
Source: [Qwen/Qwen3-8B `config.json`](https://huggingface.co/Qwen/Qwen3-8B/raw/main/config.json)
|
||||
- **Native context is 32,768 tokens**, not the 262,144 the user's brief assumed (that number belongs to
|
||||
Qwen3-4B-Instruct-2507, a different, non-reasoning 2507-refresh model — Qwen3-8B is the earlier,
|
||||
thinking-capable Qwen3 architecture with a materially smaller native window). Extending past 32K needs YaRN:
|
||||
> "Qwen3 natively supports context lengths of up to 32,768 tokens. For conversations where the total length
|
||||
> (including both input and output) significantly exceeds this limit, we recommend using RoPE scaling
|
||||
> techniques to handle long texts effectively."
|
||||
and llama.cpp-specific YaRN invocation is given explicitly:
|
||||
`./llama-cli ... -c 131072 --rope-scaling yarn --rope-scale 4 --yarn-orig-ctx 32768`, with a documented
|
||||
caveat that "all the notable open-source frameworks implement **static** YaRN, which means the scaling
|
||||
factor remains constant regardless of input length, potentially impacting performance on shorter texts."
|
||||
Source: [Qwen/Qwen3-8B-GGUF — Processing Long Texts](https://huggingface.co/Qwen/Qwen3-8B-GGUF#processing-long-texts)
|
||||
- **Weights** (official Qwen quants, fetched from the GGUF repo file list): Q5_K_M = 5.85 GB, Q8_0 = 8.71 GB.
|
||||
Source: [Qwen/Qwen3-8B-GGUF](https://huggingface.co/Qwen/Qwen3-8B-GGUF)
|
||||
*(Not independently verified: unsloth's equivalent `UD-Q4_K_XL` quant, which is what this repo's
|
||||
`docker-compose.yml`/`.env.example` actually download for every model deployed so far — the unsloth file
|
||||
size wasn't fetched, only the official Qwen quants above. Treat Q5_K_M/Q8_0 as a reasonable bound, not the
|
||||
exact file this repo would pull.)*
|
||||
- Per-token KV cache: `36 × 2 × 8 × 128 × 2 bytes = 144 KiB/token` fp16 — identical to Qwen3-4B-Instruct-2507's
|
||||
figure in `fast-model-choice.md` §5, since both share the same `layers/kv_heads/head_dim` triple.
|
||||
|
||||
| Context | KV (fp16) | KV (q8_0) | KV (q4_0, current classifier setting) |
|
||||
|---|---|---|---|
|
||||
| 65,536 (current classifier ctx) | 9.0 GiB | 4.5 GiB | **2.25 GiB** |
|
||||
| 131,072 (user's proposed "match main model") | 18.0 GiB | 9.0 GiB | **4.5 GiB** |
|
||||
|
||||
Weights + KV (q4_0, smallest realistic combo):
|
||||
|
||||
| Context | Q5_K_M weights + q4_0 KV | Q8_0 weights + q4_0 KV |
|
||||
|---|---|---|
|
||||
| 65,536 | 5.85 + 2.25 = **8.1 GB** | 8.71 + 2.25 = **10.96 GB** |
|
||||
| 131,072 | 5.85 + 4.5 = **10.35 GB** | 8.71 + 4.5 = **13.21 GB** |
|
||||
|
||||
### Qwen3-Coder-30B-A3B-Instruct
|
||||
|
||||
- Architecture (fetched from the shared Qwen3-30B-A3B-family `config.json`): `num_hidden_layers: 48`,
|
||||
`num_key_value_heads: 4`, `num_attention_heads: 32`, `head_dim: 128`, `hidden_size: 2048`, **MoE**:
|
||||
`num_experts: 128`, `num_experts_per_tok: 8` (8 of 128 experts active per token — confirms this is a sparse
|
||||
MoE model, not a dense one like the 27B or 8B candidates; the "active params" figure describes *compute*
|
||||
per token, not memory footprint — **all 128 experts' weights still have to be resident** wherever the model
|
||||
is loaded, GPU or RAM).
|
||||
Source: [Qwen/Qwen3-30B-A3B-family `config.json`](https://huggingface.co/Qwen/Qwen3-30B-A3B/raw/main/config.json)
|
||||
- **UD-Q4_K_XL file size (the exact quant/quantizer this repo already standardizes on): 17.7 GB.** 30.5B total /
|
||||
3.3B activated parameters. Native context "262,144 natively... can be extended further using Yarn to reach
|
||||
1M tokens."
|
||||
Source: [unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF](https://huggingface.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF?show_file_info=Qwen3-Coder-30B-A3B-Instruct-UD-Q4_K_XL.gguf)
|
||||
- Per-token KV cache: `48 × 2 × 4 × 128 × 2 bytes = 96 KiB/token` fp16 (smaller per-token than the 8B/4B
|
||||
candidates, since `num_key_value_heads` is 4 here vs. 8 — but this saving is irrelevant given the weights
|
||||
size below).
|
||||
|
||||
| Context | KV (fp16) | KV (q8_0) | KV (q4_0) |
|
||||
|---|---|---|---|
|
||||
| 65,536 | 6.0 GiB | 3.0 GiB | **1.5 GiB** |
|
||||
| 131,072 | 12.0 GiB | 6.0 GiB | **3.0 GiB** |
|
||||
|
||||
Weights + KV (q4_0):
|
||||
|
||||
| Context | Total |
|
||||
|---|---|
|
||||
| 65,536 | 17.7 + 1.5 = **19.2 GB** |
|
||||
| 131,072 | 17.7 + 3.0 = **20.7 GB** |
|
||||
|
||||
## 3. Does either candidate fit the ~6.1 GiB actually free on the card today?
|
||||
|
||||
**No — neither does, at either context size, even at the smallest quant/KV-quant combination tested.**
|
||||
|
||||
- Qwen3-8B's cheapest realistic combination (Q5_K_M weights + q4_0 KV at the *current* 65536 ctx, not even
|
||||
the proposed 131072) is **8.1 GB — already ~2 GB over the measured 6.1 GiB free budget**, before accounting
|
||||
for compute-buffer/batch overhead that `fast-model-choice.md` §"Implementation note" already found could add
|
||||
meaningfully on top of the naive weights+KV estimate (that's exactly why the 4B classifier ended up needing
|
||||
`--flash-attn on` and partial 28/36-layer offload instead of the originally-predicted comfortable full-GPU
|
||||
fit).
|
||||
- Qwen3-Coder-30B-A3B-Instruct isn't close at any setting tested — its weights alone (17.7 GB) are triple the
|
||||
entire free budget, and this doesn't change with context size since the weights term dominates.
|
||||
- This is a **VRAM-capacity problem, not a context-window problem** — directly confirming the "qwen8b needs to
|
||||
offload more to the RAM" intuition in the original request. Reducing context doesn't fix it; the weights
|
||||
don't fit regardless.
|
||||
|
||||
**CPU/RAM offload mechanics:** llama.cpp's `--n-gpu-layers` is documented as "max. number of layers to store in
|
||||
VRAM, either an exact number, `'auto'`, or `'all'`" — the layers not selected are computed on CPU, with the
|
||||
model loaded via mmap by default (same mechanism this repo's own `.env.example` already documents for
|
||||
`LLAMA_GPU_LAYERS`: "if GPU+RAM ever can't hold the working set, the OS pages the rest in from disk
|
||||
automatically"). For the MoE Coder-30B-A3B model specifically, this repo's own `.env.example` already flags the
|
||||
more targeted alternative — `--n-cpu-moe`/`--cpu-moe`/`--override-tensor "exps"` — as the flags that "target
|
||||
Mixture-of-Experts models (e.g. Qwen3.8-2.4T-A95B)," i.e. exactly this model's architecture: these offload only
|
||||
expert-tensor weights to CPU while keeping attention/shared layers and KV cache on GPU, which is the
|
||||
mechanically correct lever for an MoE model, unlike the blunt `--n-gpu-layers` used for the dense 8B/27B/4B
|
||||
models. **Neither llama.cpp's own README nor the Qwen model cards fetched here document a quantified
|
||||
performance cost for partial offload** — no primary source gives a "N layers offloaded = X% slower" figure.
|
||||
Source: [llama.cpp `tools/server/README.md`](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md)
|
||||
|
||||
What *is* directly measured, in this repo's own deployment history: CPU-only was tried first for the current,
|
||||
much smaller 4B classifier and rejected — "too slow in practice: real classification calls blew past
|
||||
OmniRoute's 60s timeout and retry-looped (504→499→504)" (`docker-compose.yml`'s `qwen-classifier` comment
|
||||
block). An 8B dense model has roughly double the compute of the 4B model per token; a 30B-A3B model's *routing*
|
||||
overhead on CPU (choosing 8 of 128 experts per token, each a separate weight lookup) adds a different kind of
|
||||
cost that neither this repo nor the sources fetched here have measured. **Given the classifier's Stage 1 has an
|
||||
explicit ~300ms latency budget** (`fast-model-choice.md` §1, from qwen-code's own docs), and this repo already
|
||||
has one concrete data point that CPU offload breaks that budget at a smaller model size, extending either
|
||||
candidate onto significant CPU offload carries real, unquantified latency risk — the same failure mode already
|
||||
observed once, at a favorable (smaller) model size.
|
||||
|
||||
## 4. Does a second GPU solve this, and is one actually available?
|
||||
|
||||
**Not today.** `docs/server-planing.md` is a rack-build plan for "3-4x AMD Radeon AI PRO R9700 (32GB) GPUs" —
|
||||
a future-state document, not present inventory. Every GPU-facing comment in this repo's own
|
||||
`docker-compose.yml`/`.env.example`/`rocm-gpu-pin-and-render-group.md` consistently refers to "the single 32GB
|
||||
R9700" and measures the "~6.1GiB free" budget against one physical card holding both `llama-server` and
|
||||
`qwen-classifier`. Adding a second GPU is a hardware-acquisition and rack-build question — physically sourcing,
|
||||
installing, and power/PCIe-provisioning a card per `server-planing.md`'s own build plan — not a
|
||||
`docker-compose.yml`/`.env.example` change.
|
||||
|
||||
**If a second GPU were added**, it would directly remove one already-documented risk for this specific pair:
|
||||
this repo's own `rocm-gpu-pin-and-render-group.md` traced the GPU-pinned-at-100%/ROCm#5706 bug to its precise
|
||||
trigger condition —
|
||||
|
||||
> "The pin only appears with two concurrent HIP-context-holding processes **on the same GPU**... Root cause: an
|
||||
> AMD MES (Micro Engine Scheduler) firmware bug triggered by HIP hardware-queue creation."
|
||||
Source: [ROCm/ROCm#5706](https://github.com/ROCm/ROCm/issues/5706), via this repo's own
|
||||
[`rocm-gpu-pin-and-render-group.md`](rocm-gpu-pin-and-render-group.md)
|
||||
|
||||
Since the confirmed trigger is *two HIP contexts sharing one physical card*, moving the classifier to its own,
|
||||
second GPU would put each service on a single-HIP-context card — the condition that trips the bug wouldn't
|
||||
exist for this pair anymore, and the `GPU_MAX_HW_QUEUES=1` workaround currently applied to both services
|
||||
specifically because they share one card would no longer be load-bearing for *this* pair (it would still apply
|
||||
if any future third service shared a card with either model). This wasn't independently re-verified across two
|
||||
*separate* physical cards by any source fetched in this pass — it's a direct extrapolation from the confirmed
|
||||
root cause, same category of caveat that doc's own author already flagged for its within-one-card claim.
|
||||
|
||||
## Bottom line / recommendation
|
||||
|
||||
1. **Don't apply `LLAMA_CTX_SIZE / LLAMA_PARALLEL` symmetry to the classifier.** No qwen-code primary source
|
||||
states or implies the fast/classifier model needs to match the main model's context window. The real
|
||||
requirement (§1, already established in `fast-model-choice.md`) is ~40-50K tokens worst case; the current
|
||||
`65536` already has margin. Doubling to 131072 would only double VRAM spent on KV cache for a model that
|
||||
won't otherwise fit anyway (§2-3).
|
||||
2. **Don't upgrade the classifier to Qwen3-8B or Qwen3-Coder-30B-A3B-Instruct on the current single-GPU setup.**
|
||||
Neither fits the ~6.1GiB actually free, at any context size — this is a weights-size problem, not a
|
||||
context-window problem. Forcing it would mean either (a) shrinking the main 27B model's own VRAM footprint
|
||||
to make room (a real trade-off against the primary model, not evaluated here), or (b) CPU/partial offload,
|
||||
which this repo has direct, measured evidence already breaks the classifier's latency budget at a *smaller*
|
||||
model size than either candidate.
|
||||
3. **A second GPU is the clean fix for VRAM contention and would also retire the ROCm#5706 workaround's
|
||||
relevance for this pair — but it isn't deployed hardware today.** `server-planing.md` is a future build
|
||||
plan; this is an acquisition/rack-build decision, not something achievable via a config change right now.
|
||||
4. If the actual underlying motivation is classifier *quality* (not context capacity), that's a separate,
|
||||
legitimate question this doc doesn't answer — worth its own research pass rather than solving it via a
|
||||
bigger model that doesn't fit the hardware.
|
||||
|
||||
## Sources
|
||||
|
||||
- [qwen-code docs — Configuration / Settings, `#fastmodel`](https://qwenlm.github.io/qwen-code-docs/en/users/configuration/settings/#fastmodel)
|
||||
- [qwen-code docs — Approval Mode, `#4-auto-mode---classifier-driven-approval`](https://qwenlm.github.io/qwen-code-docs/en/users/features/approval-mode/#4-auto-mode---classifier-driven-approval)
|
||||
- [qwen-code docs — Auto Mode, `#how-it-works`](https://qwenlm.github.io/qwen-code-docs/en/users/features/auto-mode/#how-it-works)
|
||||
- [Qwen/Qwen3-8B `config.json`](https://huggingface.co/Qwen/Qwen3-8B/raw/main/config.json)
|
||||
- [Qwen/Qwen3-8B-GGUF](https://huggingface.co/Qwen/Qwen3-8B-GGUF)
|
||||
- [Qwen/Qwen3-8B-GGUF — Processing Long Texts](https://huggingface.co/Qwen/Qwen3-8B-GGUF#processing-long-texts)
|
||||
- [Qwen/Qwen3-30B-A3B-family `config.json`](https://huggingface.co/Qwen/Qwen3-30B-A3B/raw/main/config.json)
|
||||
- [unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF](https://huggingface.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF?show_file_info=Qwen3-Coder-30B-A3B-Instruct-UD-Q4_K_XL.gguf)
|
||||
- [llama.cpp `tools/server/README.md`](https://raw.githubusercontent.com/ggml-org/llama.cpp/master/tools/server/README.md)
|
||||
- [ROCm/ROCm#5706](https://github.com/ROCm/ROCm/issues/5706)
|
||||
- [docs/research/fast-model-choice.md](fast-model-choice.md) (this repo — classifier transcript sizing,
|
||||
qwen-classifier's real deployment history)
|
||||
- [docs/research/qwen3.8-27b-quant.md](qwen3.8-27b-quant.md) (this repo — KV-cache-from-config.json method
|
||||
reused here)
|
||||
- [docs/research/rocm-gpu-pin-and-render-group.md](rocm-gpu-pin-and-render-group.md) (this repo — ROCm#5706
|
||||
trigger condition and `GPU_MAX_HW_QUEUES` scoping)
|
||||
- [docs/server-planing.md](../server-planing.md) (this repo — confirms only 1 of a planned 4 GPUs is deployed)
|
||||
- `docker-compose.yml`, `.env.example` (this repo — current `qwen-classifier`/`llama-server` config and the
|
||||
measured "~6.1GiB free" VRAM figure)
|
||||
|
||||
## Confidence / uncertainty summary
|
||||
|
||||
- **High confidence:** qwen-code's `fastModel`/Auto-Mode docs state no context-window requirement (direct
|
||||
quotes from all three linked pages); Qwen3-8B's native 32,768 context and YaRN caveat (direct model-card
|
||||
quote); Qwen3-Coder-30B-A3B-Instruct's MoE architecture and 17.7GB Q4_K_XL file size (direct from the
|
||||
quantizer's own repo page); the KV-cache-per-token math for both candidates (computed directly from each
|
||||
model's own `config.json`, same method already validated in this repo's prior research); the ROCm#5706
|
||||
trigger condition being scoped to two HIP contexts on the *same* GPU (direct quote from this repo's own
|
||||
prior research, itself sourced from the upstream issue).
|
||||
- **Medium confidence:** the exact unsloth `UD-Q4_K_XL`-equivalent file size for Qwen3-8B — only the official
|
||||
Qwen quants (Q5_K_M/Q8_0) were fetched, not unsloth's own repo, so the real number this repo would actually
|
||||
download wasn't directly verified (bounded reasonably by the Q5_K_M figure, which is already the smallest
|
||||
realistic option and still doesn't fit). The claim that CPU/partial-offload latency risk scales unfavorably
|
||||
for larger/MoE models is a reasoned extrapolation from this repo's one measured data point (4B CPU-only
|
||||
rejected) plus general MoE-routing-overhead reasoning, not a directly measured benchmark for either candidate.
|
||||
- **Low confidence / not independently verified:** whether `GPU_MAX_HW_QUEUES=1`/ROCm#5706 genuinely has zero
|
||||
relevance across two *separate* physical GPUs (extrapolated from the confirmed same-GPU trigger condition,
|
||||
same caveat this repo's own prior research already flagged for its own claim); no primary source found that
|
||||
quantifies llama.cpp's actual inference-speed penalty for partial `--n-gpu-layers` or `--n-cpu-moe` offload
|
||||
in general — this is a documented gap in the sources checked, not a guessed number.
|
||||
@@ -0,0 +1,222 @@
|
||||
# Evaluating Colibrì (JustVugg/colibri) for this stack
|
||||
|
||||
**Date:** 2026-09-08
|
||||
**Scope:** The user flagged https://github.com/JustVugg/colibri as something that "could revolutionize"
|
||||
this self-hosted AI stack. What is Colibrì actually, is it compatible with this stack's AMD
|
||||
ROCm/HIP-only single-GPU setup, and — even if compatible — does it fill a real gap versus what
|
||||
llama.cpp, OmniRoute, Qdrant, Neo4j, and ComfyUI already do here?
|
||||
|
||||
## 1. What Colibrì actually is
|
||||
|
||||
Colibrì is a pure-C, zero-runtime-dependency inference engine whose specific trick is treating
|
||||
"storage, RAM, and VRAM as a single inference hierarchy" so that huge mixture-of-experts (MoE)
|
||||
models — far bigger than any one machine's VRAM+RAM — can still run, by keeping the small dense
|
||||
layers resident and streaming the (much larger) set of routed experts from disk on demand with an
|
||||
LRU/"hot-store" cache and router-lookahead prefetching:
|
||||
|
||||
> "Colibrì is an open-source inference engine designed to run frontier mixture-of-experts (MoE)
|
||||
> models on consumer hardware... The fundamental approach uses 'a JIT, but for weights' — parameters
|
||||
> are staged across storage tiers (VRAM/RAM/NVMe) based on measured routing patterns rather than kept
|
||||
> resident."
|
||||
— https://raw.githubusercontent.com/JustVugg/colibri/main/README.md
|
||||
|
||||
It ships single-C-file implementations for eight specific model families — GLM-5.2/5.3,
|
||||
GLM-5.3-Flash, Inkling, Kimi K3, DeepSeek V4 Flash, Qwen3.8-Flash-Next, Qwen3.6 (35B-A3B), and OLMoE
|
||||
— each requiring model weights pre-converted into Colibrì's own container format (`coli convert`),
|
||||
not arbitrary GGUF files:
|
||||
|
||||
> "Eight model families with single C file implementations... GLM-5.2/5.3 | 744B | 372GB | 16GB+ ...
|
||||
> Kimi K3 | 2.8T | 1.6TB | 32GB+"
|
||||
— https://raw.githubusercontent.com/JustVugg/colibri/main/README.md
|
||||
|
||||
It's meant to be run either from prebuilt binaries/releases, built from source (`./setup.sh` under
|
||||
`c/`), or via Docker (`docker/Dockerfile`, `docker/Dockerfile.slim`, `docker/docker-compose.yml` exist
|
||||
in-repo — confirmed present via the GitHub contents API, https://api.github.com/repos/JustVugg/colibri/contents/docker),
|
||||
exposing an OpenAI- and Anthropic-compatible HTTP API (`coli serve`, default `http://127.0.0.1:8000/v1`,
|
||||
plus `/v1/messages`) — the same shape OmniRoute already expects from a provider, per third-party
|
||||
summaries of `docs/api.md` and `docs/serve_protocol.md`
|
||||
([search result summary, secondary](https://github.com/JustVugg/colibri/blob/main/docs/api.md)).
|
||||
|
||||
It launched July 10, 2026 and went viral on Hacker News the same day (453 points) on the strength of
|
||||
running the 744B-parameter GLM-5.2 model on a 25GB-RAM consumer box:
|
||||
|
||||
> "A new inference engine called 'Colibrì' has emerged that can run the massive AI 'GLM-5.2,' with 744
|
||||
> billion parameters, on a regular PC with 25GB of memory."
|
||||
— https://gigazine.net/gsc_news/en/20260710-colibri-glm/ (secondary coverage)
|
||||
|
||||
## 2. Hardware/runtime requirements — AMD ROCm or NVIDIA-only?
|
||||
|
||||
**This is the load-bearing question given this stack runs llama.cpp on ROCm/HIP, not CUDA, on a
|
||||
single AMD Radeon AI PRO R9700.** The answer is more nuanced than a flat yes/no — verified against
|
||||
source, not just README prose:
|
||||
|
||||
- **The engine is CPU-first; a GPU is optional at all.** `docs/quickstart.md` states plainly: "You do
|
||||
**not** need a GPU. A GPU only helps if you have one; the engine runs CPU-only by default."
|
||||
— https://raw.githubusercontent.com/JustVugg/colibri/main/docs/quickstart.md
|
||||
- **AMD/ROCm support is real, shipped, and reasonably recent — not just a README claim.** It was
|
||||
requested in [issue #69](https://github.com/JustVugg/colibri/issues/69) (opened 2026-07-11, "No
|
||||
ROCM support"), the maintainer confirmed it was mechanically straightforward since HIP closely
|
||||
mirrors CUDA, an initial PR (#112) added it but was **closed unmerged**, and a follow-up PR — tracked
|
||||
as [#339](https://github.com/JustVugg/colibri/issues/69) — landed the actual mechanism that shipped:
|
||||
a single shared CUDA kernel source (`backend_cuda.cu`) compiled either by `nvcc` or by `hipcc`
|
||||
against a compatibility header:
|
||||
|
||||
> "backend_gpu_compat.h — 'one GPU backend source, two vendors.' ... maps CUDA runtime calls to HIP
|
||||
> equivalents when compiled by hipcc with `HIP=1`... handles architecture-specific guards for rocWMMA
|
||||
> availability and matrix core support across different GPU architectures (gfx906, gfx908, gfx11xx,
|
||||
> etc.)."
|
||||
— https://raw.githubusercontent.com/JustVugg/colibri/main/c/backend_gpu_compat.h (confirmed present
|
||||
in the current `main` branch — this is not a stale/unmerged branch)
|
||||
|
||||
This shipped in a **tagged release**, not just an open PR — `CHANGELOG.md` lists "AMD GPU support" as
|
||||
part of v1.1.0 (2026-07-22): "AMD GPU support, dual-SSD streaming, fmt=5/fmt=6 quantization formats."
|
||||
— https://raw.githubusercontent.com/JustVugg/colibri/main/CHANGELOG.md
|
||||
- **Notably, the community contributor who tested it used an RX 9070 XT / gfx1201** — the same RDNA4
|
||||
architecture generation as this stack's Radeon AI PRO R9700 — per the PR #339 description: "validated
|
||||
across CPU builds, HIP testing on gfx1201, and NVIDIA compatibility verification." That's a genuinely
|
||||
favorable, non-generic signal for this specific card's GPU family, better than "AMD support exists
|
||||
somewhere."
|
||||
- **But ROCm support is thinner and less documented than the CUDA/Metal paths.** `docs/` has `cuda.md`,
|
||||
`metal.md`, `metal_implementation.md`, and `vulkan.md`, but **no `rocm.md` or `hip.md`** (confirmed via
|
||||
the GitHub contents API listing of `docs/`, https://api.github.com/repos/JustVugg/colibri/contents/docs).
|
||||
Model-specific tuning docs are written CUDA-only with no AMD mention at all — e.g.
|
||||
`docs/qwen36-cuda-tier.md` references `COLI_CUDA=1`, `backend_cuda.cu`, and lists test hardware as
|
||||
"RTX 3070 8 GB + Quadro RTX 4000" (both NVIDIA), with zero AMD/ROCm/HIP text anywhere in that
|
||||
document. — https://raw.githubusercontent.com/JustVugg/colibri/main/docs/qwen36-cuda-tier.md
|
||||
So: the *general* GPU-acceleration mechanism supports ROCm/HIP and has been community-validated on
|
||||
hardware close to the R9700, but the *per-model* tiering/tuning documentation and (presumably) most
|
||||
of the maintainer's own benchmarking is CUDA-first. Treat AMD support as functional-but-secondary,
|
||||
not a first-class, symmetrically-tested backend.
|
||||
- **A separate GPU-agnostic path also exists**: a Vulkan backend (`backend_vulkan.c`, confirmed present
|
||||
in the `c/` directory listing) that the project positions as covering "AMD via Mesa/RADV" as a
|
||||
vendor-neutral fallback, independent of the HIP path above — this would also be viable on the R9700
|
||||
in principle, though vendor-neutral compute back ends are typically slower than a vendor SDK path
|
||||
(HIP/ROCm) and no R9700/gfx1201-specific Vulkan numbers were found in the docs reviewed.
|
||||
|
||||
**Bottom line on hardware fit: not a blocker.** Unlike a hard CUDA-only dependency, Colibrì's AMD/HIP
|
||||
path is real, shipped in a release, and specifically exercised on the same RDNA4 family as this
|
||||
stack's GPU — this is not a disqualifying finding the way it would be for a CUDA-only tool.
|
||||
|
||||
## 3. License
|
||||
|
||||
Apache License 2.0, confirmed by fetching `LICENSE` directly from the repo — a standard permissive
|
||||
license, fine for self-hosted use here (commercial or non-commercial, modification, redistribution all
|
||||
permitted, patent grant included, "AS IS" with no warranty).
|
||||
— https://raw.githubusercontent.com/JustVugg/colibri/main/LICENSE
|
||||
|
||||
Model weights are licensed separately from the engine — e.g. the README notes "GLM-5.2 weights
|
||||
released by Z.ai under MIT"
|
||||
(https://raw.githubusercontent.com/JustVugg/colibri/main/README.md) — so each model's own license
|
||||
would need checking before use, same as with any GGUF today.
|
||||
|
||||
## 4. Maturity signals
|
||||
|
||||
Pulled from the GitHub API (https://api.github.com/repos/JustVugg/colibri) and the changelog:
|
||||
|
||||
| Signal | Value |
|
||||
|---|---|
|
||||
| Repo created | 2026-07-01 |
|
||||
| First tagged release (v1.0.0) | 2026-07-19 |
|
||||
| Current version (as of today) | 1.10.2 (2026-09-06) |
|
||||
| Age at time of writing | ~10 weeks |
|
||||
| Stars / Forks | 27,047 / 2,963 |
|
||||
| Open issues | 104 |
|
||||
| Top contributor | JustVugg — 1,077 commits |
|
||||
| #2 contributor | ZacharyZcR — 163 commits |
|
||||
| Total contributors | 100+ (long tail, most in single digits) |
|
||||
| License | Apache 2.0 |
|
||||
| Archived? | No |
|
||||
|
||||
Read honestly, this is **a viral, very-early-stage, single-maintainer-dominated project**, not a
|
||||
mature or slow-burn one:
|
||||
|
||||
- It is ~10 weeks old today. The star count (27k) is wildly disproportionate to that age and reflects
|
||||
a Hacker News front-page moment (453 points the day it launched,
|
||||
https://gigazine.net/gsc_news/en/20260710-colibri-glm/), not organic multi-year adoption. Stars are
|
||||
a popularity signal, not a quality proof, and here the ratio (huge stars, ~10 weeks old, one dominant
|
||||
committer) is itself a maturity red flag worth naming rather than a mark in its favor.
|
||||
- Commit/release activity is genuinely fast — 14 tagged releases in under 8 weeks (v1.0.0 on 2026-07-19
|
||||
through v1.10.2 on 2026-09-06, https://raw.githubusercontent.com/JustVugg/colibri/main/CHANGELOG.md)
|
||||
— so it is actively maintained day-to-day, not abandoned. But that pace also means breaking changes
|
||||
and security patches are frequent: v1.6.2 (2026-08-14) was itself "a security release: six
|
||||
privately-reported memory-safety issues fixed... from untrusted input," and v1.10.2 (2026-09-06)
|
||||
again touched "security fixes for image API" — both signs of a codebase still finding its footing on
|
||||
hardening, not evidence of instability being the norm, but worth weighing given this stack would be
|
||||
exposing any such server on an internal network via OmniRoute.
|
||||
- Contribution concentration is heavy: the maintainer (JustVugg) has ~6.6x the commits of the next
|
||||
contributor, and the rest of the 100+ contributor list trails off into single-digit-commit
|
||||
drive-by PRs (per the GitHub contributors API,
|
||||
https://api.github.com/repos/JustVugg/colibri/contributors) — a classic "one person's viral project
|
||||
plus a wave of small first-time PRs" shape, not an established multi-maintainer team.
|
||||
- The AMD/ROCm feature specifically has a short, thin history: requested 2026-07-11, shipped 2026-07-22
|
||||
(11 days later, in v1.1.0), documented only implicitly (no dedicated `docs/hip.md`/`rocm.md`, unlike
|
||||
every other backend) — i.e., it is the newest and least-independently-verified of the project's four
|
||||
GPU backends (CUDA, Metal, Vulkan, HIP).
|
||||
|
||||
## 5. What capability gap it would actually fill in this stack
|
||||
|
||||
Concretely comparing against what's already running (`docker-compose.yml`):
|
||||
|
||||
- **llama.cpp (ROCm) already fully GPU-resides the current model** — `--n-gpu-layers 999` on the
|
||||
llama-server service means the whole Qwen3.8-27B-class GGUF sits in the R9700's VRAM and runs at
|
||||
normal, fast, interactive token rates. Colibrì's entire value proposition is the *opposite* case:
|
||||
models **too large to fit in VRAM+RAM at all**, accepted at the cost of streaming most of the model
|
||||
from disk on every forward pass. For a model that already fits on this GPU (which is the whole point
|
||||
of the current setup), Colibrì offers no benefit — llama.cpp is already doing the fast thing.
|
||||
- **The actual gap it could fill is running models this stack categorically cannot run today** — e.g.
|
||||
GLM-5.2 (744B), DeepSeek V4 Flash (284B), or Kimi K3 (2.8T), none of which would ever fit on a single
|
||||
R9700 regardless of quantization. Colibrì's own benchmarks make the cost of that explicit and, unlike
|
||||
its capability claims, these are the project's self-reported numbers, not independently reproduced —
|
||||
flagged as such:
|
||||
|
||||
> "6× RTX 5090 (full residency): 5.8-6.8 tokens/second decode... 128GB CPU-only desktop: ~1.8
|
||||
> tokens/second (warm)... Single RTX 5070 Ti: 1.07 tokens/second... 25GB dev box: 0.05-0.1
|
||||
> tokens/second (baseline)."
|
||||
— https://raw.githubusercontent.com/JustVugg/colibri/main/README.md (self-reported, unverified by
|
||||
any third party found during this research)
|
||||
|
||||
At 0.05–2 tokens/second on any hardware remotely resembling a single-GPU workstation, this is not
|
||||
usable for the interactive coding-CLI workloads this stack is built around (Claude Code CLI, Kimi
|
||||
CLI, etc. routed through OmniRoute per the README) — those need low per-token latency for tool-calling
|
||||
round trips, not throughput measured in seconds per token. It would only be plausible as an
|
||||
occasional, patient, offline/batch capability (e.g. "let a huge model chew on a large doc overnight")
|
||||
layered in *alongside*, not instead of, the current llama.cpp path.
|
||||
- **Disk footprint is a real new cost, not a marginal one**: 167GB–1.6TB per model
|
||||
(https://raw.githubusercontent.com/JustVugg/colibri/main/README.md), which would need to be
|
||||
provisioned in addition to the existing `models` Docker volume, GGUF downloads, Qdrant/Neo4j
|
||||
volumes, and ComfyUI's model files already on this box.
|
||||
- **No overlap or replacement value for OmniRoute, Qdrant, Neo4j, or ComfyUI** — Colibrì is strictly an
|
||||
inference-engine alternative to llama.cpp for a narrow, specific list of very large MoE models; it
|
||||
does nothing related to gateway/key-management (OmniRoute's job), vector/graph storage (Qdrant/Neo4j's
|
||||
job), or image generation (ComfyUI's job). Its own `coli serve` OpenAI/Anthropic-compatible endpoint
|
||||
could in principle be registered as another OmniRoute provider the same way llama-server is today —
|
||||
that part is mechanically plausible — but it would be adding a second, much slower inference backend
|
||||
next to the existing fast one, not replacing or upgrading anything currently in the stack.
|
||||
|
||||
## 6. Bottom line
|
||||
|
||||
**Not a fit for this stack right now, and the "revolutionize" framing does not hold up** — but for a
|
||||
more specific reason than "wrong GPU vendor":
|
||||
|
||||
- **ROCm/AMD support is real and not the blocker one might expect.** It shipped in a tagged release
|
||||
(v1.1.0, 2026-07-22), lives in a compatibility header confirmed present on `main`
|
||||
(`c/backend_gpu_compat.h`), and was community-tested on an RX 9070 XT / gfx1201 — the same RDNA4
|
||||
family as this stack's R9700. This is genuinely worth noting as a positive, since it's the kind of
|
||||
thing that usually *is* disqualifying for AMD-only stacks and here it isn't.
|
||||
- **The disqualifying issue is fit, not hardware**: Colibrì solves "run a model way too big for your
|
||||
VRAM+RAM by streaming most of it from disk," at 0.05–2 tokens/second. This stack's actual situation is
|
||||
the opposite — a model sized to fully fit and run fast on a single GPU via llama.cpp. Colibrì would add
|
||||
no speed, capability, or reliability benefit to the model already running here, and its own numbers
|
||||
show it isn't fast enough to serve the interactive coding-CLI use case this stack exists for.
|
||||
- It could only ever be interesting as a *bolt-on, offline-only* capability for occasionally running an
|
||||
otherwise-impossible frontier-scale model (700B–2.8T params) for patient, non-interactive tasks — at
|
||||
the cost of hundreds of GB to ~1.6TB of extra disk per model, on a ~10-week-old, single-maintainer,
|
||||
still-hardening project (two security-patch releases already) with no ROCm-specific documentation and
|
||||
the thinnest testing history of its four GPU backends.
|
||||
- **Recommendation: worth a passing watch, not worth integrating.** Revisit if/when: (a) the project
|
||||
reaches a more established maturity point (6–12 months, broader contributor base, dedicated ROCm docs
|
||||
bringing it to parity with the CUDA/Metal paths), and (b) there's an actual concrete need in this stack
|
||||
to run a model in the 200B+ range that cannot fit on the R9700 — which isn't the case today (the
|
||||
current model is deliberately sized to fit fully in VRAM). Until then, integrating it would add
|
||||
operational surface (a new container, huge disk provisioning, a newer/less-hardened codebase) for no
|
||||
measurable improvement over the existing llama.cpp/ROCm path.
|
||||
@@ -195,6 +195,39 @@ comfortably affords the higher-precision quant.
|
||||
- [docs/research/qwen3.8-27b-tool-calling.md](qwen3.8-27b-tool-calling.md) (this repo — cross-referenced
|
||||
for the 27B model's own, still-open, tool-calling parser bugs)
|
||||
|
||||
## Implementation note (2026-09-09) — what actually shipped, and why it differs
|
||||
|
||||
The model pick (`Qwen3-4B-Instruct-2507`) held up and is what's deployed. Several sizing assumptions in
|
||||
this doc didn't survive contact with the real deployment, though — worth recording so the next person
|
||||
tuning this doesn't re-derive the same corrections from scratch:
|
||||
|
||||
- **Service name is `qwen-classifier`, not `llama-server-fast`** — this doc's proposed name never got
|
||||
used. There's no `LLAMA_FAST_CTX_SIZE`/`LLAMA_FAST_PARALLEL` in `.env.example` either; the real config
|
||||
lives inline in `docker-compose.yml`'s `qwen-classifier` command.
|
||||
- **CPU-only was tried first and rejected** — this doc's VRAM budget analysis (§5) assumed GPU
|
||||
residency from the start, but the actual rollout path tried CPU-only first (to sidestep VRAM
|
||||
contention entirely) and found it too slow: real classification calls blew past OmniRoute's request
|
||||
timeout and retry-looped. Moved to GPU after that, which is what §5's math was for all along.
|
||||
- **Q4_K_XL weights, not Q8_0** — §5's "~2.4GB headroom" case assumed Q8_0 (4.28GB). In practice, fitting
|
||||
the classifier onto the R9700 *alongside* the 27B model (not in an assumed-empty 7GB budget) left only
|
||||
~6.1GB free VRAM total, and even Q4_K_XL (2.37GB) plus full-context KV cache didn't leave enough real
|
||||
margin at full GPU offload — see the "measured live" numbers in `docker-compose.yml`'s `qwen-classifier`
|
||||
comment block. Landed on **partial GPU offload (28/36 layers)** instead of full offload, which is not a
|
||||
case this doc considered at all.
|
||||
- **65536 context, not 8192** — §5 sized the context "in the low thousands," reasoning from qwen-code's
|
||||
two-stage classifier description alone. Directly reading qwen-code's actual source
|
||||
(`packages/core/src/permissions/classifier-transcript.ts`: `MAX_TRANSCRIPT_MESSAGES=40`,
|
||||
`MAX_HISTORICAL_ACTION_CHARS=4000`/message) puts the real worst case at ~40-50K tokens — confirmed
|
||||
live, a real classifier call during testing hit 15,116 prompt tokens. 8192 would have been undersized
|
||||
for real usage; 65536 gives margin without the original setting.json value (131072, copied from the
|
||||
main model's entry, not a real qwen-code requirement) wasting VRAM for no reason.
|
||||
- **§4's `--reasoning off` recommendation was initially missed** in the first deployment pass and added
|
||||
only once this doc was re-read while writing this note. It's now in `docker-compose.yml`'s
|
||||
`qwen-classifier` command, per this doc's own "add it regardless, no-cost safety net" reasoning — still
|
||||
unconfirmed whether the current `ghcr.io/ggml-org/llama.cpp:server-rocm` build actually reproduces
|
||||
#20809 (nothing in testing so far surfaced `reasoning_content` where `tool_calls` was expected, but
|
||||
that wasn't specifically probed for either).
|
||||
|
||||
## Confidence/uncertainty summary
|
||||
|
||||
- **High confidence:** Qwen3-4B-Instruct-2507's non-thinking-only status (direct model-card quote);
|
||||
|
||||
@@ -0,0 +1,84 @@
|
||||
# OmniRoute's per-connection semaphore timeout — hardcoded, not a setting
|
||||
|
||||
**Date:** 2026-09-09
|
||||
|
||||
Any OmniRoute connection whose upstream can only handle a small, fixed number of concurrent requests
|
||||
(this repo's `llama-server`/`qwen-classifier`, both effectively single-GPU-slot-limited) can hit a hard
|
||||
30-second reject once more requests are in flight than the connection's `maxConcurrent` allows — even
|
||||
though the request would have succeeded fine if it had just waited its turn. This surfaced first as the
|
||||
`pr-agent`/`CodersPlacePI` 429/504 investigation (see the issue tracker), then again while sizing
|
||||
`qwen-classifier`. Recorded here so it doesn't have to be re-diagnosed from scratch next time.
|
||||
|
||||
## The error
|
||||
|
||||
```
|
||||
{"error":{"message":"Semaphore timeout after 30000ms for <provider>:<connectionId>","type":"rate_limit_error","code":"rate_limit_exceeded"}}
|
||||
```
|
||||
|
||||
## Root cause (confirmed against OmniRoute's own source, [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute))
|
||||
|
||||
`open-sse/services/accountSemaphore.ts`:
|
||||
|
||||
```ts
|
||||
const DEFAULT_TIMEOUT_MS = 30_000;
|
||||
...
|
||||
function createSemaphoreTimeoutError(semaphoreKey, timeoutMs) {
|
||||
const error = new Error(`Semaphore timeout after ${timeoutMs}ms for ${semaphoreKey}`);
|
||||
error.code = "SEMAPHORE_TIMEOUT"; // classified upstream as HTTP 429 rate_limit_exceeded
|
||||
return error;
|
||||
}
|
||||
```
|
||||
|
||||
Called from `open-sse/handlers/chatCore.ts`:
|
||||
|
||||
```ts
|
||||
await acquireAccountSemaphore(accountSemaphoreKey, {
|
||||
maxConcurrency: accountSemaphoreMaxConcurrency, // = the connection's maxConcurrent
|
||||
signal: streamController.signal,
|
||||
// no timeoutMs passed → always falls back to the hardcoded 30_000 default
|
||||
})
|
||||
```
|
||||
|
||||
This is **not** the same thing as OmniRoute's documented quota-share concurrency gate
|
||||
(`open-sse/services/combo/quotaShareConcurrency.ts`, key prefix `qsconn:`), which is deliberately
|
||||
fail-open per its own doc comment ("a saturated queue or timeout proceeds without a slot rather than
|
||||
ever rejecting a dispatchable request") — that one only matters for quota-share combos. The account
|
||||
semaphore above is a *different*, always-on gate keyed `provider:connectionId`, has no fail-open path,
|
||||
and its 30-second timeout is a bare `await` with nothing passed to override it — not exposed via
|
||||
`/api/resilience`, not an env var, not a dashboard toggle, not documented anywhere in
|
||||
`docs/reference/ENVIRONMENT.md`. It's a hardcoded constant in vendored code.
|
||||
|
||||
Also **not** the same as `requestQueue.maxWaitMs` (visible via `GET /api/resilience`, this deployment
|
||||
already has it at `86400000`) — that one bounds a Bottleneck-managed *execution* timer that starts only
|
||||
after dispatch, surfaces as HTTP 504 `RATE_LIMIT_EXECUTION_TIMEOUT`, and is unrelated to the 429 above.
|
||||
|
||||
## What actually fixes it
|
||||
|
||||
The 30s ceiling itself cannot be raised — no config surface reaches it in the current OmniRoute build.
|
||||
Two real options:
|
||||
|
||||
1. **Bypass the semaphore, let the upstream's own queue absorb concurrency instead.**
|
||||
`maxConcurrency == null || maxConcurrency <= 0` fully bypasses `accountSemaphore.ts` (see
|
||||
`isBypassed()`) — no gate, no 30s timer, requests pass straight through to the upstream. This only
|
||||
works if the upstream itself queues gracefully with no reject-timeout of its own — confirmed true for
|
||||
llama.cpp's server (`tools/server/server-queue.cpp` has no queue-wait timeout; excess requests just
|
||||
wait for a free slot). If you do this, also raise the connection's own
|
||||
`providerSpecificData.timeoutMs` (bounded 1ms–24h, `MAX_PROVIDER_SPECIFIC_TIMEOUT_MS`) generously —
|
||||
that's the timer that now matters: "did the upstream return response headers in time," which on
|
||||
llama.cpp means the full queue-wait-then-generate time, since llama.cpp sends **zero bytes, not even
|
||||
headers**, while a request sits queued (confirmed in `server-context.cpp`: `res->status = 200` is only
|
||||
set after the first generated token exists).
|
||||
2. **Reduce how often more than `maxConcurrent` requests actually stack up** — e.g. the
|
||||
`pr-agent`/Gitea webhook fix (narrowing the subscribed event list so one PR action doesn't fire 3+
|
||||
near-simultaneous AI calls). Doesn't remove the ceiling, just makes it less likely to be hit.
|
||||
|
||||
Applied in this repo: `llama-server`'s OmniRoute connection has `maxConcurrent: null` and
|
||||
`providerSpecificData.timeoutMs: 1200000` (20 min — matches worst-case 2-slots-busy + queued + own
|
||||
generation time). `qwen-classifier` uses a much shorter `timeoutMs: 120000` since it isn't
|
||||
GPU-contended the same way — see `docs/coding-cli-setup/qwen-code.md`.
|
||||
|
||||
## Sources
|
||||
|
||||
- [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) — `open-sse/services/accountSemaphore.ts`, `open-sse/handlers/chatCore.ts`, `open-sse/services/combo/quotaShareConcurrency.ts`, `open-sse/services/rateLimitManager.ts`, `docs/architecture/RESILIENCE_GUIDE.md`, `docs/reference/ENVIRONMENT.md`
|
||||
- [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp) — `tools/server/server-queue.cpp`, `tools/server/server-context.cpp`
|
||||
- `src/shared/validation/providerSpecificData.ts` (OmniRoute) — `MAX_PROVIDER_SPECIFIC_TIMEOUT_MS` bound
|
||||
@@ -0,0 +1,397 @@
|
||||
# A 48-minute total outage on `qwen3.8-27b-local`, and the third OmniRoute timeout mechanism this repo hadn't documented yet
|
||||
|
||||
**Date:** 2026-09-15
|
||||
|
||||
**Verdict:** The error — `"[504]: Direct response did not start within 30000ms — retrying on a fresh socket"` —
|
||||
comes from a **third, previously-undocumented OmniRoute timeout mechanism** (`OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`,
|
||||
default 30s), distinct from both timeouts already recorded in
|
||||
[`omniroute-account-semaphore-timeout.md`](./omniroute-account-semaphore-timeout.md) and
|
||||
[`omniroute-non-ping-sse-stream-timeout.md`](./omniroute-non-ping-sse-stream-timeout.md). It exists specifically to
|
||||
recover from a *stale pooled TCP socket* by retrying once on a brand-new connection — but in the incident analyzed
|
||||
here, **both** the original attempt and the fresh-socket retry timed out, repeatedly, for 46 requests over 48
|
||||
straight minutes with zero successes. That pattern rules out a stale-socket explanation (a fresh socket bypasses
|
||||
the pool entirely) and points instead at the upstream itself — `llama-server`, or the R9700 GPU underneath it —
|
||||
being genuinely unresponsive for the whole window. The best primary-source match for that symptom is an **open,
|
||||
still-unresolved AMD ROCm bug specific to this exact GPU** ([ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630)):
|
||||
an MES-firmware hang during generation on `gfx1201`/R9700 that leaves the process alive but stuck, sometimes for
|
||||
no logged reason at all. No config change fixes this — raising the 30s timeout only makes each failed attempt
|
||||
take longer to give up, it doesn't un-wedge a hung GPU.
|
||||
|
||||
## The evidence
|
||||
|
||||
**Source:** `omniroute-request-logs-6h-2026-09-15.json`, a 339-entry OmniRoute request-log export the user pulled
|
||||
from the dashboard, covering `2026-09-15T12:49:52Z`–`18:44:54Z`. Every entry with a non-200 status (66 of them)
|
||||
is a `POST /v1/chat/completions` against `qwen3.8-27b-local` (`/models/Qwen3.8-27B-UD-Q4_K_XL.gguf`), all on the
|
||||
same `connectionId` (`649a2d3e-7527-488e-9b8a-dc4ac2624176`) and the same `provider`
|
||||
(`openai-compatible-chat-a7bda643-6687-41f4-b75d-fd2cab746874`) — a single upstream connection, not a fan-out
|
||||
artifact.
|
||||
|
||||
Sorting every error by timestamp shows **one continuous outage**, not scattered slow requests:
|
||||
|
||||
- Last successful `/v1/chat/completions` before the outage: `15:18:12.285Z`
|
||||
- **First failure:** `15:25:06.415Z` — status 504, `"[504]: Direct response did not start within 30000ms —
|
||||
retrying on a fresh socket"`, duration `60186ms`
|
||||
- **Every single `/v1/chat/completions` attempt** from `15:25:06.415Z` through `16:13:29.502Z` failed — 46× 504
|
||||
(all clustered `60021`-`60324ms`, i.e. two back-to-back 30s attempts, both failing) interleaved with 20× 499
|
||||
(`"Request aborted"` / `"Client disconnected: request_signal_aborted"`, durations `1.4s`-`99.97s` — these are
|
||||
qwen-code giving up client-side while OmniRoute was still mid-retry)
|
||||
- **First successful recovery:** `16:13:50.388Z`, `20873ms` — 21 minutes after the last failure attempt cluster,
|
||||
i.e. the very next attempt after the outage window succeeded normally
|
||||
- No successful `/v1/chat/completions` call appears anywhere inside the `15:25:06Z`-`16:13:29Z` window — confirmed
|
||||
by filtering all 154 `/v1/chat/completions` log entries in that range: every one is 504 or 499.
|
||||
|
||||
`git log --since=2026-09-14 --until=2026-09-16` shows **zero commits** in this repo on 2026-09-15 — the outage
|
||||
correlates with no deploy, `scripts/update.sh` run, or config change on this end.
|
||||
|
||||
**A second, independent data point** (see caveat below): the user also pasted a large raw text table, copied
|
||||
directly from OmniRoute's dashboard UI rather than the JSON export, showing the **`qwen-classifier`** connection
|
||||
(`qwen3-4b`, both the `-UD-Q4_K_XL.gguf` file this repo's `docker-compose.yml` currently defaults to, and a
|
||||
`-UD-Q8_K_XL.gguf` variant that appears **nowhere** in this repo's checked-in `docker-compose.yml`/`.env.example`
|
||||
— either tested by hand against `LLAMA_CLASSIFIER_MODEL_FILE` outside version control, or evidence of drift worth
|
||||
checking directly on the server) failing repeatedly across two accounts (`Haylan`, `qwen-cli-main`), with the
|
||||
same signature: `TI: 0|TO: 0` (zero tokens either direction — failed before generating anything) and durations
|
||||
clustering at exactly `60.0`-`60.4s` for the 504s. That's the identical two-attempts-at-30s-each shape as the 27B
|
||||
outage above, strongly suggesting the same underlying mechanism, though — important caveat — **this table is not
|
||||
present in the 6-hour JSON export and covers a different, longer time range** (timestamps back to `01:21` and
|
||||
`23:54` on unspecified dates), so it cannot be directly time-correlated against the 27B outage above. Treat it as
|
||||
corroborating evidence that this failure mode recurs on both local model connections, not as proof they failed at
|
||||
the same moment.
|
||||
|
||||
## Root cause: `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`, a mechanism built for a different problem
|
||||
|
||||
Confirmed directly against OmniRoute's own source, [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute)
|
||||
(same repo the two existing timeout docs already cite) — `open-sse/utils/directResponseStartTimeout.ts`:
|
||||
|
||||
```ts
|
||||
const DEFAULT_DIRECT_HEADERS_TIMEOUT_MS = 30_000;
|
||||
const DIRECT_RESPONSE_START_TIMEOUT_CODE = "DIRECT_RESPONSE_START_TIMEOUT";
|
||||
|
||||
export function resolveDirectHeadersTimeoutMs(
|
||||
env: Record<string, string | undefined> = process.env
|
||||
): number {
|
||||
const raw = env.OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS;
|
||||
if (raw == null || raw.trim() === "") return DEFAULT_DIRECT_HEADERS_TIMEOUT_MS;
|
||||
...
|
||||
}
|
||||
|
||||
function createDirectResponseStartTimeout(timeoutMs: number): Error & { code: string } {
|
||||
const err = new Error(
|
||||
`Direct response did not start within ${timeoutMs}ms — retrying on a fresh socket`
|
||||
) as Error & { code: string };
|
||||
...
|
||||
}
|
||||
```
|
||||
|
||||
"Direct" here means **direct (no-proxy) egress** — confirmed in `open-sse/utils/proxyFetch.ts`, which routes any
|
||||
connection with no configured upstream HTTP proxy through this path (as opposed to OmniRoute's separate
|
||||
proxy/relay egress paths). Every local connection in this repo (`llama-server`, `qwen-classifier`, both reached
|
||||
over the `ai-stack` Docker network with no proxy) is "direct" — so this timeout mechanism governs **every**
|
||||
request to either local model, streaming or non-streaming alike, not just non-streaming JSON responses as the
|
||||
name might suggest.
|
||||
|
||||
### Why it exists (and why it didn't help here)
|
||||
|
||||
Two OmniRoute issues, both with dedicated regression tests in the repo, explain the actual design intent:
|
||||
|
||||
- **#4252** (`tests/unit/proxyfetch-retry-fresh-socket-4252.test.ts`): "Undici dispatcher fails on direct provider
|
||||
requests in 502 bursts" — the default direct dispatcher pools keep-alive sockets; some upstreams silently close
|
||||
idle pooled sockets, so the next request reusing one fails with `UND_ERR_SOCKET`. Fix: retry once on a **fresh,
|
||||
no-keep-alive dispatcher** (`getRetryDispatcher()`, a different instance from `getDefaultDispatcher()`) so the
|
||||
retry can't grab another already-dead pooled socket.
|
||||
- **#10214** (`tests/unit/proxyfetch-direct-response-start-timeout-10214.test.ts`): "Direct (no-proxy) requests
|
||||
stall on a silently-dropped pooled keep-alive socket until the caller's deadline or a service restart" — the
|
||||
harder case: a pooled socket that dies **without even an error**, just silence. Undici's `headersTimeout`
|
||||
default (600s) is far too slow to catch this in practice, and the existing #4252 retry never fires because no
|
||||
error is thrown to trigger it. The fix bounds each direct attempt's response-start wait to
|
||||
`OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS` (30s default) via `directFetchWithBoundedResponseStart`, and retries
|
||||
**once** on the same fresh no-keep-alive dispatcher from #4252 when that bound is hit.
|
||||
|
||||
Both fixes assume the *socket* is the problem, not the upstream. `directFetchWithBoundedResponseStart`'s own
|
||||
implementation (`open-sse/utils/directResponseStartTimeout.ts`) is exactly two attempts: pooled, then fresh. When
|
||||
attempt 2 — a brand-new socket that cannot possibly be a zombie pooled connection — **also** times out at 30s,
|
||||
the retry logic has nothing left to try and the request fails with `DIRECT_RESPONSE_START_TIMEOUT_CODE`
|
||||
(surfaced as the 504 seen in the logs). A fresh socket succeeding to *connect* but the *server* never sending a
|
||||
response is exactly what "the upstream process is alive but stuck" looks like from OmniRoute's side — it can't
|
||||
distinguish "GPU is wedged mid-generation" from "stale pooled socket," because both present as "nothing came
|
||||
back in 30s." 46 consecutive both-attempts-failed cycles over 48 minutes is far outside what a transient stale-socket
|
||||
burst (the scenario #4252/#10214 were built for) would produce; it's consistent with a sustained upstream
|
||||
outage instead.
|
||||
|
||||
**Not currently configured in this repo**: `grep`-ing `docker-compose.yml` and `.env.example` for
|
||||
`OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS` finds nothing — this deployment runs on the unmodified 30s default. (This
|
||||
is also, notably, a *fourth* data point alongside the two already-documented mechanisms and the `requestQueue.maxWaitMs`/
|
||||
`RATE_LIMIT_EXECUTION_TIMEOUT` timer mentioned in passing in `omniroute-account-semaphore-timeout.md` — OmniRoute
|
||||
has at least four independent timeout knobs guarding different stages of a request's life, three of them
|
||||
30-second-flavored by default, which is worth keeping in mind the next time an unfamiliar timeout string shows up.)
|
||||
|
||||
## Why the upstream itself was likely unresponsive: ROCm/legacy-rocm-build#6630
|
||||
|
||||
This session has **no SSH/shell access to the actual R9700 server** — there's no SSH config, and nothing in
|
||||
`scripts/` does remote exec, confirmed by inspecting `scripts/update.sh` and the absence of any `~/.ssh/config`
|
||||
entry for the box. So none of the following is confirmed against this specific incident's `dmesg`/`rocm-smi`/
|
||||
`docker logs` output — it's the closest primary-source match to the *symptom*, not a diagnosis of *this* outage.
|
||||
|
||||
[ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630) ("gfx1201 R9700 ROCm 7.14
|
||||
llama.cpp generation hang, MES queue failure and PSP reset -62; Vulkan passes") is an **open**, actively
|
||||
investigated issue (created 2026-08-19, most recent update 2026-08-28, no fix landed) that reproduces on **the
|
||||
exact same GPU this repo runs on** — Radeon AI PRO R9700, `gfx1201` — running **llama.cpp with `-fa on`** (this
|
||||
repo's `llama-server` also runs `--flash-attn on`), with a controlled Vulkan-vs-ROCm A/B: Vulkan completes
|
||||
normally, ROCm hangs during token generation. Direct quotes:
|
||||
|
||||
> "During the ROCm generation stall: GPU busy reached 100%, memory busy remained 0%, VRAM use was only about 1.1
|
||||
> GB, **the container remained alive but made no output progress**."
|
||||
|
||||
> "when MES stops responding, **the driver can stay unaware of it indefinitely** — the failure only surfaces when
|
||||
> something happens to send the next MES message... if a run is left alone after it stops making progress, the
|
||||
> kernel prints nothing at all, so a hang can look like a slow workload rather than a fault."
|
||||
|
||||
> "the failure is probabilistic, not deterministic... a single passing run on this host does not indicate a
|
||||
> healthy configuration."
|
||||
|
||||
The thread (12 comments as of this research pass, an AMD engineer `harkgill-amd` participating) has ruled out,
|
||||
one at a time, `uni_mes=0`, `mes_log_enable=1`, ROCm 6.4.4, ROCm 7.14, ROCm 10.0.0 stable, the latest TheRock
|
||||
nightly, and GFXOFF-disable — **no confirmed fix or workaround exists in the thread as of this research pass**.
|
||||
A related comment on the same issue (`chrisfranson`) reports the identical MES `REMOVE_QUEUE`/MODE1-reset
|
||||
signature from a **completely unrelated workload** (headless LibreOffice with OpenCL) on the same `gfx1201`
|
||||
silicon, reinforcing that this is a driver/firmware-level fault under general GPU load, not something specific
|
||||
to llama.cpp's request pattern.
|
||||
|
||||
**This is a different bug from the one already mitigated in this repo.** `docs/research/rocm-gpu-pin-and-render-group.md`
|
||||
already documents and works around `ROCm/ROCm#5706` (clock/power pinned at boost whenever two concurrent HIP
|
||||
contexts share the GPU — fixed via `GPU_MAX_HW_QUEUES=1`, already set on both `llama-server` and `qwen-classifier`
|
||||
in `docker-compose.yml`). #5706's symptom is elevated power draw with the GPU still working; #6630's symptom is
|
||||
generation fully halting with `gpu_busy=100%`/`mem_busy=0%` and MES no longer responding at all — a real hang, not
|
||||
a clock-pin inefficiency. `GPU_MAX_HW_QUEUES=1` targets #5706's specific trigger (hardware-queue oversubscription
|
||||
across concurrent HIP processes) and has no evidence in #6630's thread of affecting that bug — #6630 reproduces
|
||||
in single-GPU, single-process benchmarks with no second HIP context involved at all, so the already-applied fix
|
||||
should not be assumed to help here.
|
||||
|
||||
## What would actually resolve this vs. what wouldn't
|
||||
|
||||
- **Raising `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`** — not recommended as a fix. It would make each failed attempt
|
||||
take longer before giving up (worse latency during a real hang), without addressing why the GPU stopped
|
||||
responding. It's the right lever only if future evidence shows genuinely-slow-but-working responses being
|
||||
mistaken for hangs (the same shape as the already-fixed `REQUEST_TIMEOUT_MS` issue in
|
||||
`omniroute-non-ping-sse-stream-timeout.md`) — this incident's 46-for-46 both-attempts-failed pattern over 48
|
||||
minutes doesn't fit that shape.
|
||||
- **Concrete next step for whoever has server access when this recurs**: check `dmesg | grep -i amdgpu` and
|
||||
`journalctl -k` on the R9700 host for `MES(...) failed to respond`, `GPU reset begin`, or `PSP resume failed`
|
||||
lines matching #6630's signature, and `docker logs llama-server`/`docker logs qwen-classifier` to see whether
|
||||
the process was alive-but-stuck (consistent with #6630) versus crashed/restarted (which would point elsewhere).
|
||||
Capturing this during a live incident is the only way to move this from "best primary-source match" to
|
||||
"confirmed root cause."
|
||||
- **Monitor [ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630)** for a fix —
|
||||
it's open and active (AMD engineer engaged as of 2026-08-26); no released ROCm version as of this research pass
|
||||
is confirmed clean.
|
||||
- **The `-UD-Q8_K_XL.gguf` classifier variant in the pasted table but absent from version control** is worth a
|
||||
direct look on the server (`cat .env` / `docker inspect qwen-classifier` for the actual `LLAMA_CLASSIFIER_MODEL_FILE`
|
||||
in effect) — outside this research pass's reach without server access, flagged here so it isn't lost.
|
||||
|
||||
## Recovery: how the hang actually clears (or doesn't) — addendum, 2026-09-15
|
||||
|
||||
Follow-up question: what actually recovers the socket once this hits, given it's been observed to stay wedged
|
||||
for days at a time? Pulled the full comment thread on
|
||||
[ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630) directly via the GitHub API
|
||||
(12 comments, `angelhalo` as primary reporter, `harkgill-amd` as the responding AMD engineer, plus one
|
||||
corroborating report from `chrisfranson` on unrelated hardware/workload) — the earlier research pass's source list
|
||||
cited this issue but hadn't read the full thread. Three things fall directly out of it:
|
||||
|
||||
**There is no reliable in-band recovery.** The driver doesn't notice the hang on its own — direct quote:
|
||||
"when MES stops responding, the driver can stay unaware of it indefinitely — the failure only surfaces when
|
||||
something happens to send the next MES message." In a captured live hang, "every driver-managed ring is
|
||||
completely idle while the GPU reports 100% busy," `dmesg` has zero amdgpu lines, and no task is in D-state —
|
||||
so from the OS's perspective nothing is wrong; only sending the GPU another command (which killing/restarting
|
||||
the stuck process does) triggers the driver to discover the wedge and attempt its own MODE1 reset.
|
||||
|
||||
**Once triggered, that reset itself is a coin flip across three documented outcomes**, not a guaranteed fix:
|
||||
1. **Clean recovery** — `GPU reset succeeded, trying to resume`, PSP resumes, the card comes back (VRAM is
|
||||
wiped — "VRAM is lost due to GPU reset!" — so the container needs a real restart to reload the model, a plain
|
||||
process respawn isn't enough even when the reset itself works).
|
||||
2. **Failed resume** — `PSP resume failed`, `GPU reset end with ret = -62` (the original report's own outcome) —
|
||||
the reset attempt itself fails, leaving the GPU in a worse state than before.
|
||||
3. **Full kernel soft-lockup** — `chrisfranson`'s independent report (different workload — headless LibreOffice
|
||||
OpenCL, different card — RX 9070 XT, same `gfx1201` silicon) hit outcome 2 or 3 twice out of three times:
|
||||
"the whole system hard-locked (kernel soft lockup pegging a CPU at ~90-100% softirq, requiring a physical
|
||||
power cycle)."
|
||||
|
||||
`angelhalo` deliberately left one hang untouched rather than killing the process, to observe it without
|
||||
contaminating the state with a reset: "I recovered only with a subsequent cold power cycle" — no reset was ever
|
||||
triggered because nothing sent the GPU another message. Their standard test procedure between every single run
|
||||
in this thread is "a cold power cycle (AC removed, ≥30 s)," specifically **not** a warm/soft reboot — stated
|
||||
reason: "on this card a MODE1 reset takes the host down with it," meaning even the OS's own reboot path can't
|
||||
be trusted to come back cleanly once this GPU is in a bad state. This is the practical answer to "why does it
|
||||
stay stuck for days": nothing about the hang self-clears, `docker`'s `restart: unless-stopped` policy never
|
||||
fires because the container process is alive and never exits (confirmed: this repo's `llama-server` and
|
||||
`qwen-classifier` services have no `healthcheck` block at all — only `omniroute` itself does, a plain TCP
|
||||
connect check on its own dashboard port, which says nothing about whether `llama-server`/`qwen-classifier` are
|
||||
responding) — so a hang persists until a human notices the symptom (requests failing) and manually intervenes,
|
||||
and "days" is just however long that takes to notice on a homelab box, not a property of the hang itself.
|
||||
|
||||
**No fix or reliable mitigation exists as of this reading (2026-08-28, the thread's latest comment).**
|
||||
`harkgill-amd` (AMD) could not reproduce locally and asked for a nightly-driver retest; `angelhalo` retested and
|
||||
it still failed. Every other variable tested still hangs: ROCm 6.4.4 through 10.0.0 stable, TheRock nightlies,
|
||||
`amdgpu.uni_mes=0`, `cwsr_enable=0`, `mes_log_enable=1`, GFXOFF disabled, two different physical R9700 cards, both
|
||||
llama.cpp and vLLM. The thread's own conclusion, as of the last comment: "a probabilistic lost-completion event"
|
||||
with no known trigger to avoid and no known driver/firmware combination that's clean.
|
||||
|
||||
**Practical takeaway for this repo, given no upstream fix exists:**
|
||||
- A restart *might* recover it, *might* make it worse (failed PSP resume), and *might* take the whole host down
|
||||
requiring a physical power cycle — there's no way to know in advance which outcome a given hang will produce.
|
||||
- Nothing currently watches for this automatically. Docker's `restart: unless-stopped` is the wrong tool (process
|
||||
doesn't exit) — recovering automatically would need a `healthcheck` against `llama-server`'s own `/health`
|
||||
endpoint (llama.cpp's built-in liveness endpoint) paired with something that acts on an `unhealthy` status,
|
||||
since Docker itself doesn't restart on failed healthchecks without an external watcher (e.g. `willfarrell/autoheal`
|
||||
or equivalent) — not evaluated here, flagged as a real gap, not a recommendation to implement blind: an
|
||||
automated restart during a hang that's about to fail its PSP resume and lock the host could turn a
|
||||
"requests are failing" incident into "the box needs a physical power cycle" automatically and unattended,
|
||||
which is a real downside worth weighing against faster detection.
|
||||
- Given the reset outcome is unpredictable, the safest manual recovery when this is caught live is: restart the
|
||||
affected container, then immediately check `dmesg | grep -i amdgpu` for `PSP resume failed` or a soft-lockup
|
||||
signature before assuming it's fixed — if either appears, a full reboot (and per this thread's own testing
|
||||
practice, possibly a genuine AC power cycle rather than a warm reboot) is the next step, not a second restart
|
||||
attempt.
|
||||
|
||||
## Caveats and open questions
|
||||
|
||||
- **JSON export vs. pasted table are two different, non-overlapping captures.** The JSON file is a precise 6-hour
|
||||
window with full per-request detail; the pasted table is a longer, dashboard-UI-copied range with less
|
||||
structure and no verifiable overlap with the JSON file's timestamps. A fresh multi-day JSON export (same
|
||||
`request-logs` endpoint used to produce the file analyzed here) would let a future pass check whether the
|
||||
classifier's failures and the 27B model's outage are literally simultaneous (strong evidence for a shared
|
||||
GPU-level cause) or independent recurrences of the same mechanism on separate schedules.
|
||||
- **Why the outage self-recovered after ~48 minutes with no observed restart is unexplained.** #6630's thread
|
||||
describes hangs resolving via an explicit GPU reset (sometimes failing, requiring reboot) — not a case of a
|
||||
hang clearing on its own after a fixed interval. Nothing in the available data (no server access) confirms
|
||||
whether a restart happened that isn't visible from OmniRoute's logs, or whether this specific hang genuinely
|
||||
self-cleared, which would be a data point *against* the #6630 hypothesis worth capturing next time.
|
||||
- **Live reproduction was not attempted.** The user suggested testing tool-calls against the classifier via the
|
||||
Windows-side qwen-code CLI (`C:\Users\aerli\AppData\Local\qwen-code\bin\qwen.cmd`) to try to reproduce a
|
||||
"Direct response did not start" failure live. Skipped for this pass: qwen-code requires an interactive/
|
||||
already-authenticated session to drive meaningfully, and deliberately trying to reproduce a GPU hang against
|
||||
the shared production classifier risked a genuine 60s+ stall on infrastructure other work depends on, for
|
||||
uncertain diagnostic payoff given the strength of the log-based and source-based evidence already gathered.
|
||||
Worth doing deliberately, with server access on hand to capture `rocm-smi`/`dmesg` simultaneously, rather than
|
||||
as a quick check from this pass.
|
||||
|
||||
## Sources
|
||||
|
||||
- `omniroute-request-logs-6h-2026-09-15.json` — OmniRoute dashboard request-log export provided by the user
|
||||
(2026-09-15, 339 entries, `12:49:52Z`-`18:44:54Z`)
|
||||
- User-pasted OmniRoute dashboard table (`qwen-classifier`/`qwen3-4b` failures, separate capture window)
|
||||
- [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) —
|
||||
`open-sse/utils/directResponseStartTimeout.ts`, `open-sse/utils/proxyFetch.ts`, `open-sse/utils/proxyDispatcher.ts`,
|
||||
`tests/unit/proxyfetch-direct-response-start-timeout-10214.test.ts`, `tests/unit/proxyfetch-retry-fresh-socket-4252.test.ts`,
|
||||
`open-sse/handlers/chatCore/upstreamTimeouts.ts`
|
||||
- [ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630) — open R9700/gfx1201
|
||||
llama.cpp generation-hang issue, full comment thread (2026-08-19 through 2026-08-28)
|
||||
- [`docs/research/omniroute-account-semaphore-timeout.md`](./omniroute-account-semaphore-timeout.md) — the
|
||||
first already-documented 30s OmniRoute timeout (account semaphore, 429, hardcoded)
|
||||
- [`docs/research/omniroute-non-ping-sse-stream-timeout.md`](./omniroute-non-ping-sse-stream-timeout.md) — the
|
||||
second already-documented timeout (first-SSE-event deadline, `REQUEST_TIMEOUT_MS`-derived)
|
||||
- [`docs/research/rocm-gpu-pin-and-render-group.md`](./rocm-gpu-pin-and-render-group.md) — the already-mitigated,
|
||||
*different* R9700/gfx1201 MES bug ([ROCm/ROCm#5706](https://github.com/ROCm/ROCm/issues/5706), clock-pin/power,
|
||||
not a hang)
|
||||
- Local `docker-compose.yml`, `.env.example` (grepped directly, confirming `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`
|
||||
is unset / on the 30s default, and that `GPU_MAX_HW_QUEUES=1` is already applied to both GPU services)
|
||||
- `git log --since=2026-09-14 --until=2026-09-16` (this repo, confirming zero commits during the outage window)
|
||||
|
||||
## Confidence / uncertainty summary
|
||||
|
||||
- **High confidence**: the exact mechanism and semantics of `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS` /
|
||||
`directFetchWithBoundedResponseStart` (read directly from OmniRoute's own source and its two regression test
|
||||
files, which spell out the intent in comments referencing the originating issues); the JSON log's chronology,
|
||||
single-connection scope, and the "two ~30s attempts, both failing" shape of every 504 (computed directly from
|
||||
the log file); that this timeout is unconfigured in this repo (direct grep); that no commits landed in this
|
||||
repo during the outage window (direct `git log`).
|
||||
- **Medium confidence**: that ROCm/legacy-rocm-build#6630 is the actual root cause of *this specific* outage.
|
||||
The GPU model, driver-family symptom shape (`gpu_busy=100%`/`mem_busy=0%`, alive-but-stuck, sometimes
|
||||
logged/sometimes silent), and `-fa on` usage all match closely, and it's an open/unresolved/actively-discussed
|
||||
issue as of this research pass — but nothing from this incident's own `dmesg`/`rocm-smi` output was available
|
||||
to confirm it directly (no SSH access), so this is the best primary-source match to the symptom, not a
|
||||
confirmed diagnosis.
|
||||
- **Low confidence / open**: why the outage recovered on its own after ~48 minutes with no observed restart;
|
||||
whether the classifier's pasted-table failures share the exact same triggering event as the 27B outage
|
||||
analyzed here (same failure signature, but no verified time-overlap between the two data sources); the
|
||||
provenance of the `-UD-Q8_K_XL.gguf` classifier variant seen in the pasted table but absent from version
|
||||
control.
|
||||
|
||||
## Live test, 2026-09-15: the classifier is measurably too slow at its own documented worst case — independent of any hang
|
||||
|
||||
Before scoping a fix, tested the live `qwen-classifier` backend directly against realistic worst-case load,
|
||||
per the user's request to gather fresh evidence rather than design blind. Two attempts to reproduce this through
|
||||
qwen-code itself first surfaced an unrelated, separately-useful finding; the direct backend test below is what
|
||||
actually answered the question.
|
||||
|
||||
### qwen-code's own headless mode never reaches the classifier
|
||||
|
||||
Ran `qwen --approval-mode auto <prompt>` (positional/one-shot, non-interactive) from the Windows-side install
|
||||
(`C:\Users\aerli\AppData\Local\qwen-code\bin\qwen.cmd`, which has both `fastModel` and the `omniroute-search` MCP
|
||||
server already configured), asking it to run a shell `dir` and use the web-search MCP tool. Both attempts hit a
|
||||
wall before any classifier request was even sent:
|
||||
|
||||
- The MCP tool call was refused outright: `Warning: Tool "mcp__omniroute-search__search" requires user approval
|
||||
but cannot execute in non-interactive mode. ... use the -y flag (YOLO mode)`.
|
||||
- The shell tool: the model itself reported `run_shell_command` as "not registered" in this session and silently
|
||||
substituted a read-only `glob` call instead — no approval prompt, no classifier call, no system warning printed
|
||||
(unlike the MCP case), across two separate clean runs.
|
||||
|
||||
**Conclusion: one-shot headless `qwen <prompt>` invocations don't exercise Auto Mode's classifier at all for
|
||||
approval-requiring tools** — they're declined or silently rerouted before the classifier ever gets a request.
|
||||
The classifier only fires in a genuinely interactive session, where it substitutes for the human's live approval
|
||||
decision. This wasn't previously documented anywhere in this repo and is worth keeping in mind: headless qwen-code
|
||||
testing is not a valid way to probe classifier behavior, live or otherwise. (Not investigated further: whether
|
||||
`qwen serve`/`--input-format stream-json` headless-agent modes behave differently — plausible, since they're
|
||||
built for exactly this kind of automation, but out of scope for this pass.)
|
||||
|
||||
### Direct backend test: real classifier latency at realistic token counts
|
||||
|
||||
Given headless qwen-code couldn't drive this, sent shaped classifier requests straight to
|
||||
`http://proxy-ai.home/v1/chat/completions` (model `qwen3-4b//models/Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`,
|
||||
confirmed present in `GET /v1/models`) using a temporary scoped API key, mimicking the `{shouldBlock}`
|
||||
JSON-verdict shape and the `MAX_TRANSCRIPT_MESSAGES=40` / `MAX_HISTORICAL_ACTION_CHARS=4000` structure this
|
||||
repo's own `fast-model-choice.md` already read out of qwen-code's `classifier-transcript.ts` source. Five calls,
|
||||
in order:
|
||||
|
||||
| # | Prompt tokens | `cached_tokens` | Wall time | Notes |
|
||||
|---|---|---|---|---|
|
||||
| 1 | 1,000 | 3 | 2.24s | Small prompt, genuinely fresh — healthy baseline. |
|
||||
| 2 | 51,234 | 51,233 | **64.21s** | First send of a large synthetic worst-case transcript (~40 highly self-similar 4K-char "historical action" blocks). Near-total cache hit reported, yet still the slowest call — see caveat below. |
|
||||
| 3 | 51,234 | 51,233 | 0.07s | **Identical repeat of #2.** Same `chatcmpl-...` id as #2 came back — this is OmniRoute short-circuiting an exact-duplicate request via a proxy-level response cache, not fresh inference. Confirms #2/#3's `cached_tokens` field is not a reliable proxy for wall-clock latency on its own. |
|
||||
| 4 | 1,000 | 3 | 0.03s | Identical repeat of #1 — same id, same response-cache short-circuit. |
|
||||
| 5 | 30,958 | 919 | **28.66s** | Fresh, non-repeated worst-case-shaped transcript (different random content, no internal self-similarity to trigger cache effects). Mostly-uncached (919/30,958) — this is the clean data point. |
|
||||
|
||||
Call #5 is the one to trust: **~31K genuinely-fresh prompt tokens took 28.7 seconds** on the current
|
||||
`--n-gpu-layers 28` (of 36) / `--cache-type-k/v q4_0` / `--parallel 1` configuration, with the backend otherwise
|
||||
idle and healthy (no hang in progress). Extrapolating that rate to the repo's own documented worst case (a real
|
||||
15,116-token call observed live per `fast-model-choice.md`, and a theoretical ceiling around 40-50K tokens per
|
||||
`classifier-transcript.ts`'s limits) puts a genuine worst-case classifier call at **roughly 30-65 seconds of
|
||||
normal, non-hung processing time** — consistent with call #2's 64.21s, even though that call's own cache
|
||||
metadata is too muddied by internal prompt self-similarity to use as a second clean sample.
|
||||
|
||||
**This directly overlaps both binding timeouts**: qwen-code's own client-side classifier stage timeout
|
||||
(`stage1Ms`/`stage2Ms`, `60000` each in the Windows-side `settings.json` observed this session, `30000`/`60000`
|
||||
in the WSL-side one) and `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`'s 30-second-per-attempt window documented above.
|
||||
A worst-case classifier call landing in the 30-65s range will **routinely** trip one or both of these timeouts
|
||||
on its own, with the backend never having hung at all — the exact same 499/504/"Request aborted" shape as the
|
||||
GPU-hang hypothesis produces, but from an entirely mundane, deterministic cause: **the classifier's current
|
||||
configuration is simply too slow for the request sizes qwen-code's own classifier-transcript design allows.**
|
||||
|
||||
### What this changes
|
||||
|
||||
This doesn't rule out ROCm/legacy-rocm-build#6630 — the 27B model's 48-minute total outage (zero successes, not
|
||||
just slow ones) doesn't fit a "just slow" explanation, and remains best matched by a real GPU hang. But it does
|
||||
mean **the classifier's own recurring failures (the pasted-table evidence) very plausibly have a second,
|
||||
independent, non-probabilistic cause that a healthcheck/restart-on-hang design wouldn't fix at all** — restarting
|
||||
a classifier that's merely slow-but-working at worst-case load just interrupts a call that would have succeeded,
|
||||
and would fire repeatedly under normal peak usage, not just during a rare hang. Any fix that only targets "detect
|
||||
and recover from an unresponsive GPU" leaves this second failure mode untouched. Two independent levers worth
|
||||
weighing before finalizing a scope: raising the classifier's own timeouts to match its real worst-case latency
|
||||
(cheap, immediate, but does nothing for actual hangs), and/or speeding up the classifier itself (full GPU offload
|
||||
if VRAM allows, a faster quant, or capping the transcript size client-side) to bring worst-case latency back
|
||||
under the existing timeouts.
|
||||
|
||||
**Not investigated in this pass**: whether call #2's 64.21s (vs. call #5's extrapolated ~45-48s at a similar
|
||||
token count) reflects genuine non-linear slowdown at the very largest context sizes, real concurrent contention
|
||||
from other production traffic sharing the same `--parallel 1` slot during the test, or is just noise from a
|
||||
single sample each — worth a few more clean, uniquely-content, worst-case-sized calls at different times of day
|
||||
before treating either number as precise.
|
||||
@@ -0,0 +1,216 @@
|
||||
# OmniRoute's builtin memory tools silently hijack qwen-code's classifier tool-call, not a model or GPU problem
|
||||
|
||||
**Date:** 2026-09-15
|
||||
|
||||
**Verdict:** The `"Classifier stage 1 unavailable"` / `"Auto Mode couldn't classify this action"` failures are **not**
|
||||
a GPU hang, not a timeout, and not a Qwen3-4B quality problem. Confirmed directly from a live debug log: the fast
|
||||
model (`qwen3-4b//models/Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`) *is* being used (`stage=fast` in every classifier
|
||||
log line), and it responds well within its timeout (18.7s against a 30-60s budget). The failure is
|
||||
`"Error: Invalid side query response: params must have required property 'shouldBlock'"` — a **schema-validation
|
||||
failure on an in-time response**. Root cause, confirmed directly against OmniRoute's own source
|
||||
(`open-sse/handlers/chatCore/memorySkillsInjection.ts`): **OmniRoute silently appends its own builtin memory tools
|
||||
(`memory_save`/`update`/`search`/`delete`) to every non-streaming chat completion's `tools` array**, whenever
|
||||
memory is enabled for the calling key — regardless of what tools the caller declared. qwen-code's classifier forces
|
||||
`tool_choice: ANY` (call *some* tool, not a specific one) so it can get a structured `respond_in_schema` JSON
|
||||
response. With OmniRoute's extra tools spliced in, the small model sometimes picks the injected `memory_save` tool
|
||||
instead — and since qwen-code's client only extracts the classifier's answer from a `respond_in_schema` function
|
||||
call (not from a stray `memory_save` call, even if the answer also happens to be present as plain text), the
|
||||
result validated against `STAGE1_SCHEMA` is empty, producing exactly the observed error.
|
||||
|
||||
## The debug-log evidence
|
||||
|
||||
Captured directly from a `-d` (debug) qwen-code run, `C:\Users\aerli\.qwen\debug\886d00eb-...txt`:
|
||||
|
||||
```
|
||||
21:09:56 [DEBUG] [CLASSIFIER] ALLOW stage=fast tool=mcp__omniroute-search__search durationMs=15412
|
||||
21:10:28 [WARN] [CLASSIFIER] failUnavailable stage=fast durationMs=18727 reason="Classifier stage 1 unavailable" cause="Error: Invalid side query response: params must have required property 'shouldBlock'"
|
||||
```
|
||||
|
||||
Both lines are tagged `stage=fast` — qwen-code's own internal label confirming the classifier used the configured
|
||||
fast model both times, settling a live question this session raised about whether the classifier was silently
|
||||
falling back to the main 27B model. It wasn't. The second call's 18.7s duration is comfortably inside this
|
||||
deployment's configured `stage1Ms` (30000-60000ms depending on install) — this was not a timeout.
|
||||
|
||||
## Reproduced directly
|
||||
|
||||
Sent the exact same request shape qwen-code's `classifyAction()`/`generateJson()` builds — a forced tool call to a
|
||||
synthetic `respond_in_schema` function with `STAGE1_SCHEMA` (`{required: ["shouldBlock"], properties: {shouldBlock:
|
||||
{type: "boolean"}}}`) as its only declared tool, `tool_choice: "required"` — directly against
|
||||
`POST /v1/chat/completions` on this deployment (model `qwen3-4b//models/Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`).
|
||||
The very first response:
|
||||
|
||||
```json
|
||||
{
|
||||
"message": {
|
||||
"content": "{\"shouldBlock\": true}",
|
||||
"tool_calls": [{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "memory_save",
|
||||
"arguments": "{\"key\": \"tool_call_review\", \"content\": \"web_fetch call to https://www.swfte.com/ai/leaderboard was blocked due to potential data exfiltration or unauthorized external access.\", \"type\": \"factual\"}"
|
||||
}
|
||||
}]
|
||||
},
|
||||
"tool_results": [{ "output": "{\"success\":true,\"memory\":{\"id\":\"84389721-...\", ...}, \"message\":\"Memory saved successfully\"}" }]
|
||||
}
|
||||
```
|
||||
|
||||
`memory_save` was never declared in the request's `tools` array — only `respond_in_schema` was. OmniRoute added it
|
||||
server-side, the model called it instead of (or alongside) the intended tool, and OmniRoute **auto-executed it**
|
||||
(`tool_results` shows a real memory record was created, with a UUID and timestamp — this isn't inert). The correct
|
||||
`{"shouldBlock": true}` answer is sitting right there in `content`, but qwen-code's `generateJson()` only ever
|
||||
looks for it inside a `respond_in_schema` function call (`functionCalls.find(call => call.name ===
|
||||
"respond_in_schema")`); when that's absent it falls through to a raw-text JSON parse of `getResponseText(result)`
|
||||
— and OpenAI-client conventions generally don't populate a response's plain-text field when `finish_reason` is
|
||||
`tool_calls` rather than `stop`, so that fallback likely returns nothing, yielding `{}` against a schema that
|
||||
requires `shouldBlock`. This is a single confirmed repro, not a guaranteed-every-time reproduction — the injected
|
||||
tool doesn't win every time (an immediate identical follow-up request returned an OmniRoute-cached copy of the same
|
||||
response, not a fresh sample — see the cache caveat in
|
||||
[`omniroute-direct-response-timeout-outage-2026-09-15.md`](./omniroute-direct-response-timeout-outage-2026-09-15.md)),
|
||||
but it reproduces the *exact* failure shape from the live debug log on the first genuine attempt.
|
||||
|
||||
## Root cause, confirmed in OmniRoute's own source
|
||||
|
||||
[diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) —
|
||||
`open-sse/handlers/chatCore/memorySkillsInjection.ts`:
|
||||
|
||||
```ts
|
||||
if (memoryOwnerId && memorySettings?.enabled && body.stream !== true) {
|
||||
// Server-side builtin memory tools (memory_save/update/search/delete) are
|
||||
// executed by the gateway's tool-call interception, which runs only on the
|
||||
// non-stream path. Stream clients (opencode etc.) execute tools client-side,
|
||||
// so for them these tools would be announced but never executed; they should
|
||||
// use the MCP memory tools (omniroute_memory_*) instead.
|
||||
const existingTools = Array.isArray(body.tools) ? body.tools : [];
|
||||
...
|
||||
const memoryTools = buildMemoryToolsForProvider(...).filter(tool => !existingToolNames.has(name));
|
||||
if (memoryTools.length > 0) {
|
||||
body = { ...body, tools: [...existingTools, ...memoryTools] };
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This runs unconditionally for any non-streaming request from a key with memory enabled — there is no exemption
|
||||
for a caller that already set `tool_choice` to force a *specific* tool. qwen-code's classifier is exactly this
|
||||
case: a single-purpose, forced-`ANY`, non-streaming tool call, which is precisely the shape this injection logic
|
||||
was not written to avoid interfering with.
|
||||
|
||||
`src/lib/memory/settings.ts` confirms `enabled: false` is the *default* — memory is off by default in a fresh
|
||||
OmniRoute install specifically because of injected-context cost, per its own comment:
|
||||
|
||||
> "Off by default: enabling memory injects up to `maxTokens` (~2k) of retrieved context into every chat request,
|
||||
> which is billed — a surprising cost for new installs... Opt in explicitly via Settings → Memory... Per-request
|
||||
> opt-out is also available via the `x-omniroute-no-memory` header."
|
||||
|
||||
This deployment has memory enabled (confirmed live by the reproduction above), which is presumably a deliberate
|
||||
choice for other workflows (chat memory across sessions) — but it has an undocumented-to-this-repo side effect on
|
||||
any caller using forced-tool-call classification.
|
||||
|
||||
## What would fix this
|
||||
|
||||
Two per-target exclusions were checked live against this deployment and confirmed **not to exist**:
|
||||
|
||||
- **Per-API-key memory override**: `GET /api/keys` was fetched directly (the temporary key handed to this session
|
||||
turned out to carry admin access, well beyond the plain `/v1` workload scope its name implied). Every key's full
|
||||
field list was inspected — `noLog`, `scopes`, `allowedModels`, `rateLimits`, `disableNonPublicModels`, etc. — with
|
||||
no memory-related field anywhere.
|
||||
- **Per-model/connection override**: `GET /api/providers/<id>` for the classifier's own connection
|
||||
(`qwen3-4b`, id `b78ceb4c-52f8-47ae-b245-483baa6e3fc2`) was fetched directly. `providerSpecificData` (`prefix`,
|
||||
`apiType`, `baseUrl`, `nodeName`, `timeoutMs`, `apiKeyHealth`) has no memory field either — consistent with the
|
||||
source: `memoryOwnerId` is resolved purely from the *calling key* (`resolveMemoryOwnerId(apiKeyInfo)`), before
|
||||
OmniRoute has even picked a provider, so it can't know or care that this particular request targets the
|
||||
classifier model specifically.
|
||||
|
||||
**The fix that was actually available and is now applied**: `x-omniroute-no-memory`, OmniRoute's own per-*request*
|
||||
opt-out (not per-key or per-model), confirmed end-to-end and traced through both sides:
|
||||
|
||||
- OmniRoute's handling, confirmed directly in `open-sse/handlers/chatCore.ts` and its own test suite
|
||||
(`tests/unit/no-memory-header.test.ts`): `memoryOwnerId = isNoMemoryRequested(headers) ? null : resolveMemoryOwnerId(...)`
|
||||
— a null owner id short-circuits *both* branches in `injectMemoryAndSkills` (context injection and tool
|
||||
injection). The test suite gives the exact accepted values: `"true"`, `"1"`, `"yes"` (case-insensitive on both
|
||||
the header name and value); `"false"`/`"0"`/`"no"`/empty do not trigger it.
|
||||
- qwen-code's support for sending it, confirmed against the installed bundle, *not* just the docs: `modelProviders.
|
||||
openai[].generationConfig.customHeaders` (documented at
|
||||
[model-providers](https://qwenlm.github.io/qwen-code-docs/en/users/configuration/model-providers/)) flows into
|
||||
`DefaultOpenAICompatibleProvider.buildClient()` (`chunk-CXTPVBFA.js`), which passes it straight into the
|
||||
underlying `OpenAI` SDK client as `defaultHeaders` — applied to every request made through that one model entry,
|
||||
and *only* that entry (confirmed this is the generic OpenAI-compatible-chat client, the same one the classifier's
|
||||
`apiType: "chat"` connection uses — not Anthropic- or Responses-API-specific plumbing).
|
||||
|
||||
**Applied**, 2026-09-15: added to the `qwen3-4b-classifier` entry in the Windows-side `~/.qwen/settings.json`
|
||||
(`C:\Users\aerli\.qwen\settings.json`), inside its `generationConfig`, alongside the existing `contextWindowSize`
|
||||
and `extra_body`:
|
||||
|
||||
```json
|
||||
"customHeaders": { "x-omniroute-no-memory": "true" }
|
||||
```
|
||||
|
||||
Scoped to this one model entry only — the main `qwen3.8-27b-local` connection's `generationConfig` is untouched,
|
||||
so its own memory-context behavior (if any is relied on elsewhere) is unaffected. qwen-code's own
|
||||
`[MODEL_PROVIDERS_HOT_RELOAD]` settings watcher (confirmed present in this session's debug log) should pick this
|
||||
up on the already-running session without a restart. **Not yet verified live** — the next classifier failure (or a
|
||||
deliberate repro, per the "Reproduced directly" section above) should confirm no `memory_save`-shaped tool call
|
||||
appears in the response once this is in effect.
|
||||
|
||||
- **Remaining fallback, if the header approach doesn't hold up**: disable memory globally for this deployment
|
||||
(`PATCH /api/settings/memory`, `enabled: false`, or Settings → Memory in the dashboard) — blunt, but confirmed to
|
||||
work by definition since `enabled: false` is every fresh install's default.
|
||||
- **Also worth doing regardless**: file this upstream with OmniRoute. Their own code already special-cases one
|
||||
caller type (streaming clients) right next to this injection logic; a similar exemption for a caller that already
|
||||
set `tool_choice` to force one specific tool would be a clean fix on their end that doesn't depend on every
|
||||
client remembering to send an opt-out header.
|
||||
- **Not a fix, and not the problem**: nothing on the classifier-model or llama.cpp side. Qwen3-4B-Instruct-2507
|
||||
correctly produced the right answer (`{"shouldBlock": true}`) in the one reproduction captured here — the model
|
||||
was never at fault.
|
||||
|
||||
## Scope note
|
||||
|
||||
This session's earlier hypothesis that the 27B model's 48-minute total outage
|
||||
([`omniroute-direct-response-timeout-outage-2026-09-15.md`](./omniroute-direct-response-timeout-outage-2026-09-15.md))
|
||||
was caused by a ROCm/gfx1201 GPU hang is set aside here per explicit direction, not retracted — that was a
|
||||
different incident (zero successes for 48 straight minutes, a shape this memory-injection bug doesn't produce) and
|
||||
this finding doesn't bear on it either way.
|
||||
|
||||
## Sources
|
||||
|
||||
- Live debug log, `C:\Users\aerli\.qwen\debug\886d00eb-5b2b-4d84-b1ef-60909f75eec2.txt` (this session, 2026-09-15)
|
||||
- Direct reproduction against this deployment's `POST /v1/chat/completions` (this session, 2026-09-15)
|
||||
- Live `GET /api/keys`, `GET /api/providers`, `GET /api/providers/b78ceb4c-52f8-47ae-b245-483baa6e3fc2`,
|
||||
`GET /api/settings/memory` against this deployment's OmniRoute instance (this session, 2026-09-15) — confirmed no
|
||||
per-key or per-connection memory field exists in either schema
|
||||
- [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) —
|
||||
`open-sse/handlers/chatCore/memorySkillsInjection.ts`, `open-sse/handlers/chatCore.ts` (the
|
||||
`isNoMemoryRequested`/`resolveMemoryOwnerId` branch), `src/lib/memory/settings.ts`, `src/lib/memory/injection.ts`,
|
||||
`open-sse/mcp-server/tools/memoryTools.ts`, `tests/unit/no-memory-header.test.ts` (exact accepted header
|
||||
name/value set)
|
||||
- [Qwen Code docs — Model Providers](https://qwenlm.github.io/qwen-code-docs/en/users/configuration/model-providers/)
|
||||
(`customHeaders` field, documented under `generationConfig`)
|
||||
- Installed qwen-code bundle — `chunk-N7VWZDWW.js`, `chunk-HBU7EKY4.js` (`classifyAction`, `runSideQuery`,
|
||||
`resolveDefaultModel`, `generateJson`, `resolveFastModelSelector`, `getFastModel`) and `chunk-CXTPVBFA.js`
|
||||
(`DefaultOpenAICompatibleProvider.buildHeaders()`/`buildClient()`, confirming `customHeaders` reaches the actual
|
||||
OpenAI SDK client as `defaultHeaders` for the plain chat-completions path the classifier uses) — all read
|
||||
directly from the bundled (unminified variable names) source, not inferred from docs alone
|
||||
- Applied fix: `C:\Users\aerli\.qwen\settings.json`, `qwen3-4b-classifier` entry's `generationConfig.customHeaders`
|
||||
(this session, 2026-09-15)
|
||||
|
||||
## Confidence / uncertainty summary
|
||||
|
||||
- **High confidence**: the fast model is genuinely used for classification (`stage=fast` in qwen-code's own debug
|
||||
log, both on success and failure); the failure is a schema-validation error on an in-time response, not a
|
||||
timeout (18.7s duration, explicit error text); OmniRoute's `memorySkillsInjection.ts` unconditionally injects
|
||||
builtin memory tools into non-streaming completions for any memory-enabled key, with no exemption for
|
||||
forced-single-tool callers (read directly from source); no per-key or per-model/connection memory override
|
||||
exists in this OmniRoute version (confirmed by reading the complete live schema of both, not by absence of
|
||||
documentation); `x-omniroute-no-memory: true` is a real, working per-request opt-out on OmniRoute's side (its
|
||||
own test suite) and is reachable from qwen-code via `modelProviders.openai[].generationConfig.customHeaders`,
|
||||
traced to the exact HTTP client the classifier's connection type uses (not inferred from docs alone — confirmed
|
||||
against the bundled source's actual header-merging code).
|
||||
- **Medium confidence**: that this exact tool-injection mechanism explains the *specific* production failures seen
|
||||
earlier in this session's testing (the reproduction matches the failure shape and the source confirms the
|
||||
mechanism exists and applies to this call pattern, but the live debug-log failure itself wasn't captured
|
||||
mid-flight with response inspection — only its aftermath, the error message).
|
||||
- **Low confidence / not verified**: the exact conditions under which the model picks the injected tool over the
|
||||
intended one (one clean reproduction on the first attempt, not a characterized hit rate — the failure may not be
|
||||
deterministic, so the `customHeaders` fix should still be watched rather than assumed to have fully resolved it
|
||||
on the strength of this write-up alone); whether the applied `customHeaders` fix has been confirmed live yet
|
||||
(not as of this writing — see "Applied" above).
|
||||
@@ -0,0 +1,60 @@
|
||||
# OmniRoute's "non-ping SSE" first-token deadline — a different timer than `STREAM_IDLE_TIMEOUT_MS`
|
||||
|
||||
**Date:** 2026-09-10
|
||||
|
||||
`STREAM_IDLE_TIMEOUT_MS` was raised to 180000 on 2026-09-09 (see docker-compose.yml's `omniroute`
|
||||
service) specifically to give contended `llama-server` prefill room to produce a first token. It didn't
|
||||
work: the very next morning, qwen-code sessions against `qwen3.8-27b-local` still hit repeated
|
||||
|
||||
```
|
||||
Stream produced no non-ping SSE event within 95000ms
|
||||
```
|
||||
|
||||
(and once at 115000ms) — both well under the 180s the compose fix set, and well under the connection's
|
||||
own `providerSpecificData.timeoutMs: 1200000` (confirmed live via `GET /api/providers/<id>`). Neither of
|
||||
those settings bounds this failure.
|
||||
|
||||
## Root cause
|
||||
|
||||
Per OmniRoute's own docs (`docs/reference/ENVIRONMENT.md`, "Timeout Settings" section) and a maintainer
|
||||
reply in [diegosouzapw/OmniRoute#10602](https://github.com/diegosouzapw/OmniRoute/discussions/10602):
|
||||
|
||||
| Variable | Default | Governs |
|
||||
|---|---|---|
|
||||
| `REQUEST_TIMEOUT_MS` | 600000 (10 min) | Overall upstream request budget. **The first non-ping SSE event's deadline inherits this one.** |
|
||||
| `STREAM_IDLE_TIMEOUT_MS` | 120000 (2 min) | Max gap between *successive* SSE chunks once streaming has already started — does not govern the wait for the first chunk. |
|
||||
| `STREAM_PING_INTERVAL_MS` | 30000 (30s) | How often OmniRoute emits its own keepalive pings on the stream — these explicitly do not count as "non-ping" events, so they can't rescue a request against the first deadline. |
|
||||
|
||||
So the 2026-09-09 fix tuned the wrong timer for this failure mode: `STREAM_IDLE_TIMEOUT_MS` only matters
|
||||
once `llama-server` has already emitted something. The "no token at all yet" case — exactly what a large
|
||||
compact-prompt prefill on a contended local model produces — is bounded by `REQUEST_TIMEOUT_MS` instead.
|
||||
|
||||
The observed 95000ms/115000ms figures are also *not* `REQUEST_TIMEOUT_MS`'s raw 600000ms default: OmniRoute
|
||||
computes the first-event deadline as **remaining budget**, not a flat timer — `REQUEST_TIMEOUT_MS` minus
|
||||
time already spent in OmniRoute's own request-queue/retry/cooldown cycle (`requestRetry: 3`,
|
||||
`connectionCooldown.apikey.baseCooldownMs`, provider breaker) before the request was actually dispatched
|
||||
to `llama-server`. Confirmed live via `GET /api/settings` → `resilienceSettings` on this deployment. Most
|
||||
of the 10-minute default budget was being burned by retries before the final attempt even started.
|
||||
|
||||
## Fix
|
||||
|
||||
Set `REQUEST_TIMEOUT_MS` explicitly, generously — applied in docker-compose.yml as
|
||||
`OMNIROUTE_REQUEST_TIMEOUT_MS` (default 1800000 / 30 min), same pattern as
|
||||
`OMNIROUTE_STREAM_IDLE_TIMEOUT_MS`. This doesn't replace the 2026-09-09 `STREAM_IDLE_TIMEOUT_MS` fix —
|
||||
that one still matters for mid-stream stalls after generation has started — it addresses the separate
|
||||
"nothing has arrived yet" case that fix didn't cover.
|
||||
|
||||
Raising `REQUEST_TIMEOUT_MS` buys headroom; it doesn't address *why* prefill on a 50K+ token compact
|
||||
prompt can take that long in the first place. `llama-server` had no `--cache-reuse` flag set — every
|
||||
request reprefilled its full prompt from scratch even when most of a conversation's prefix was unchanged
|
||||
from the previous turn. Added `--cache-reuse 256` (docker-compose.yml) so llama.cpp reuses cached KV for
|
||||
any matching ≥256-token chunk via KV-shift instead of reprocessing it, which is the actual fix for
|
||||
compact-prompt prefill time — the timeout bump above is a safety margin around it, not a substitute.
|
||||
|
||||
## Sources
|
||||
|
||||
- [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) — `docs/reference/ENVIRONMENT.md`
|
||||
("Timeout Settings"), [Discussion #10602](https://github.com/diegosouzapw/OmniRoute/discussions/10602)
|
||||
- Live `GET /api/providers/<connectionId>`, `GET /api/settings`, `GET /api/resilience` against this
|
||||
deployment's OmniRoute instance (2026-09-10)
|
||||
- [`docs/research/omniroute-account-semaphore-timeout.md`](./omniroute-account-semaphore-timeout.md) — the related-but-distinct 30s semaphore/429 investigation
|
||||
@@ -0,0 +1,64 @@
|
||||
# Ponytail audit — repo-wide over-engineering scan, 2026-09-09
|
||||
|
||||
Whole-tree audit (ponytail-audit skill), not a diff review. The only real
|
||||
code in this repo is `scripts/update.sh` (253L) and `scripts/switch-model.sh`
|
||||
(50L), plus `docker-compose.yml` and `.env.example`; the rest is docs.
|
||||
`switch-model.sh` and the compose comments (ROCm GID workarounds,
|
||||
`GPU_MAX_HW_QUEUES` rationale, lazytainer label placement) are load-bearing
|
||||
and lean — left alone. Scope: over-engineering and complexity only;
|
||||
correctness/security/performance out of scope. Findings ranked biggest cut
|
||||
first. One-shot report — nothing was applied.
|
||||
|
||||
## Findings
|
||||
|
||||
1. **`delete:` the `downloader-fast` line — it references a service removed
|
||||
in `5d6a17f` ("feat: remove llama-server-fast") and no longer exists in
|
||||
`docker-compose.yml`.** Under `set -euo pipefail`,
|
||||
`docker compose run downloader-fast` errors on the unknown service and
|
||||
**aborts every `update.sh` run** right after config sync, before omniroute
|
||||
comes up. Dead code that also breaks the mandatory deploy flow.
|
||||
*Remove the line.* [scripts/update.sh:235]
|
||||
|
||||
2. **`delete:` the entire gum path — `ensure_gum()` (~27L, L59–85), the
|
||||
`GUM_VERSION`/`GUM_DIR`/`GUM_BIN` vars (L56–58), the vendored
|
||||
`scripts/vendor/gum_0.14.5_Linux_x86_64.tar.gz` (4.4MB checked into git),
|
||||
the download fallback, and the `if [ -n "$gum_bin" ]` branch (L123,
|
||||
L127–130).** The plain-bash fallback (L131–148) already makes the
|
||||
*identical* decision (which keys take the new value) whenever gum is
|
||||
absent; the gum TUI is a speculative nicer prompt on top of a working
|
||||
path. ~4.4MB in git + arch detection + a `.cache/gum` layer, all to
|
||||
prettify a rare interactive conflict. *Replacement: nothing — always use
|
||||
the plain-bash one-screen prompt.* [scripts/update.sh, scripts/vendor/]
|
||||
|
||||
3. **`delete:` stale `llama-server-fast` / `fastModel` references — the fast
|
||||
model was removed but the docs still describe a two-model Qwen Code
|
||||
setup.** `.env.example:98` comment still lists it;
|
||||
`docs/coding-cli-setup/index.md:32` says "2 models: chat + `fastModel`";
|
||||
and `docs/coding-cli-setup/qwen-code.md` carries a whole fast-model
|
||||
section (11 refs: the `fastModel` config block, `LLAMA_FAST_CTX_SIZE`
|
||||
notes, Qwen3-4B). *Rewrite to single-model.* [docs/coding-cli-setup/
|
||||
qwen-code.md, index.md, .env.example:98]
|
||||
|
||||
4. **`shrink:` the 3× repeated `test -f … || curl …` blocks in
|
||||
`downloader-comfyui` (YAML L85–99, ~15 lines) → a `for` loop over the 3
|
||||
model files (~5 lines).** *Low confidence:* the env var names are
|
||||
non-uniform (`COMFYUI_DIFFUSION_MODEL_FILE` / `TEXT_ENCODER_FILE` /
|
||||
`VAE_FILE`), so the loop needs a small `case` — marginal win, and it
|
||||
matches the house "one-off downloader" style. [docker-compose.yml]
|
||||
|
||||
5. **`yagni:` (verify-first) qdrant + neo4j run with no consumer in the
|
||||
stack yet** — added ahead of the RAG app via the `feat-rag-databases`
|
||||
merge; nothing writes to them. Two always-on DBs for a feature that
|
||||
isn't wired. *Confirm the RAG consumer is still on the roadmap before
|
||||
keeping both; the compose comment already concedes neo4j "can absorb
|
||||
qdrant's job later."* Low confidence — deliberate tracked decision, and
|
||||
cheap to leave running. [docker-compose.yml]
|
||||
|
||||
## Net
|
||||
|
||||
`net: -45 lines script/compose (+~15 stale doc lines), -1 dep (gum, 4.4MB
|
||||
vendored binary) possible.`
|
||||
|
||||
No out-of-scope (correctness/security/performance) findings. #1 is the one
|
||||
to fix first — it's not just bloat, it's the deploy script halting on every
|
||||
run.
|
||||
@@ -233,6 +233,7 @@ docker compose build --pull
|
||||
echo "==> ensuring models are downloaded (skips already-present files)"
|
||||
docker compose --profile tools run --rm downloader
|
||||
docker compose --profile tools run --rm downloader-fast
|
||||
docker compose --profile tools run --rm downloader-classifier
|
||||
docker compose --profile tools run --rm downloader-comfyui
|
||||
|
||||
echo "==> bringing up omniroute"
|
||||
|
||||
Reference in New Issue
Block a user