docs(research): trace 48-min outage to third OmniRoute timeout and open ROCm#6630 GPU hang

This commit is contained in:
2026-09-15 23:31:58 +02:00
parent fa852c8fb4
commit ed90256f43
@@ -0,0 +1,397 @@
# A 48-minute total outage on `qwen3.8-27b-local`, and the third OmniRoute timeout mechanism this repo hadn't documented yet
**Date:** 2026-09-15
**Verdict:** The error — `"[504]: Direct response did not start within 30000ms — retrying on a fresh socket"`
comes from a **third, previously-undocumented OmniRoute timeout mechanism** (`OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`,
default 30s), distinct from both timeouts already recorded in
[`omniroute-account-semaphore-timeout.md`](./omniroute-account-semaphore-timeout.md) and
[`omniroute-non-ping-sse-stream-timeout.md`](./omniroute-non-ping-sse-stream-timeout.md). It exists specifically to
recover from a *stale pooled TCP socket* by retrying once on a brand-new connection — but in the incident analyzed
here, **both** the original attempt and the fresh-socket retry timed out, repeatedly, for 46 requests over 48
straight minutes with zero successes. That pattern rules out a stale-socket explanation (a fresh socket bypasses
the pool entirely) and points instead at the upstream itself — `llama-server`, or the R9700 GPU underneath it —
being genuinely unresponsive for the whole window. The best primary-source match for that symptom is an **open,
still-unresolved AMD ROCm bug specific to this exact GPU** ([ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630)):
an MES-firmware hang during generation on `gfx1201`/R9700 that leaves the process alive but stuck, sometimes for
no logged reason at all. No config change fixes this — raising the 30s timeout only makes each failed attempt
take longer to give up, it doesn't un-wedge a hung GPU.
## The evidence
**Source:** `omniroute-request-logs-6h-2026-09-15.json`, a 339-entry OmniRoute request-log export the user pulled
from the dashboard, covering `2026-09-15T12:49:52Z``18:44:54Z`. Every entry with a non-200 status (66 of them)
is a `POST /v1/chat/completions` against `qwen3.8-27b-local` (`/models/Qwen3.8-27B-UD-Q4_K_XL.gguf`), all on the
same `connectionId` (`649a2d3e-7527-488e-9b8a-dc4ac2624176`) and the same `provider`
(`openai-compatible-chat-a7bda643-6687-41f4-b75d-fd2cab746874`) — a single upstream connection, not a fan-out
artifact.
Sorting every error by timestamp shows **one continuous outage**, not scattered slow requests:
- Last successful `/v1/chat/completions` before the outage: `15:18:12.285Z`
- **First failure:** `15:25:06.415Z` — status 504, `"[504]: Direct response did not start within 30000ms —
retrying on a fresh socket"`, duration `60186ms`
- **Every single `/v1/chat/completions` attempt** from `15:25:06.415Z` through `16:13:29.502Z` failed — 46× 504
(all clustered `60021`-`60324ms`, i.e. two back-to-back 30s attempts, both failing) interleaved with 20× 499
(`"Request aborted"` / `"Client disconnected: request_signal_aborted"`, durations `1.4s`-`99.97s` — these are
qwen-code giving up client-side while OmniRoute was still mid-retry)
- **First successful recovery:** `16:13:50.388Z`, `20873ms` — 21 minutes after the last failure attempt cluster,
i.e. the very next attempt after the outage window succeeded normally
- No successful `/v1/chat/completions` call appears anywhere inside the `15:25:06Z`-`16:13:29Z` window — confirmed
by filtering all 154 `/v1/chat/completions` log entries in that range: every one is 504 or 499.
`git log --since=2026-09-14 --until=2026-09-16` shows **zero commits** in this repo on 2026-09-15 — the outage
correlates with no deploy, `scripts/update.sh` run, or config change on this end.
**A second, independent data point** (see caveat below): the user also pasted a large raw text table, copied
directly from OmniRoute's dashboard UI rather than the JSON export, showing the **`qwen-classifier`** connection
(`qwen3-4b`, both the `-UD-Q4_K_XL.gguf` file this repo's `docker-compose.yml` currently defaults to, and a
`-UD-Q8_K_XL.gguf` variant that appears **nowhere** in this repo's checked-in `docker-compose.yml`/`.env.example`
— either tested by hand against `LLAMA_CLASSIFIER_MODEL_FILE` outside version control, or evidence of drift worth
checking directly on the server) failing repeatedly across two accounts (`Haylan`, `qwen-cli-main`), with the
same signature: `TI: 0|TO: 0` (zero tokens either direction — failed before generating anything) and durations
clustering at exactly `60.0`-`60.4s` for the 504s. That's the identical two-attempts-at-30s-each shape as the 27B
outage above, strongly suggesting the same underlying mechanism, though — important caveat — **this table is not
present in the 6-hour JSON export and covers a different, longer time range** (timestamps back to `01:21` and
`23:54` on unspecified dates), so it cannot be directly time-correlated against the 27B outage above. Treat it as
corroborating evidence that this failure mode recurs on both local model connections, not as proof they failed at
the same moment.
## Root cause: `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`, a mechanism built for a different problem
Confirmed directly against OmniRoute's own source, [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute)
(same repo the two existing timeout docs already cite) — `open-sse/utils/directResponseStartTimeout.ts`:
```ts
const DEFAULT_DIRECT_HEADERS_TIMEOUT_MS = 30_000;
const DIRECT_RESPONSE_START_TIMEOUT_CODE = "DIRECT_RESPONSE_START_TIMEOUT";
export function resolveDirectHeadersTimeoutMs(
env: Record<string, string | undefined> = process.env
): number {
const raw = env.OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS;
if (raw == null || raw.trim() === "") return DEFAULT_DIRECT_HEADERS_TIMEOUT_MS;
...
}
function createDirectResponseStartTimeout(timeoutMs: number): Error & { code: string } {
const err = new Error(
`Direct response did not start within ${timeoutMs}ms — retrying on a fresh socket`
) as Error & { code: string };
...
}
```
"Direct" here means **direct (no-proxy) egress** — confirmed in `open-sse/utils/proxyFetch.ts`, which routes any
connection with no configured upstream HTTP proxy through this path (as opposed to OmniRoute's separate
proxy/relay egress paths). Every local connection in this repo (`llama-server`, `qwen-classifier`, both reached
over the `ai-stack` Docker network with no proxy) is "direct" — so this timeout mechanism governs **every**
request to either local model, streaming or non-streaming alike, not just non-streaming JSON responses as the
name might suggest.
### Why it exists (and why it didn't help here)
Two OmniRoute issues, both with dedicated regression tests in the repo, explain the actual design intent:
- **#4252** (`tests/unit/proxyfetch-retry-fresh-socket-4252.test.ts`): "Undici dispatcher fails on direct provider
requests in 502 bursts" — the default direct dispatcher pools keep-alive sockets; some upstreams silently close
idle pooled sockets, so the next request reusing one fails with `UND_ERR_SOCKET`. Fix: retry once on a **fresh,
no-keep-alive dispatcher** (`getRetryDispatcher()`, a different instance from `getDefaultDispatcher()`) so the
retry can't grab another already-dead pooled socket.
- **#10214** (`tests/unit/proxyfetch-direct-response-start-timeout-10214.test.ts`): "Direct (no-proxy) requests
stall on a silently-dropped pooled keep-alive socket until the caller's deadline or a service restart" — the
harder case: a pooled socket that dies **without even an error**, just silence. Undici's `headersTimeout`
default (600s) is far too slow to catch this in practice, and the existing #4252 retry never fires because no
error is thrown to trigger it. The fix bounds each direct attempt's response-start wait to
`OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS` (30s default) via `directFetchWithBoundedResponseStart`, and retries
**once** on the same fresh no-keep-alive dispatcher from #4252 when that bound is hit.
Both fixes assume the *socket* is the problem, not the upstream. `directFetchWithBoundedResponseStart`'s own
implementation (`open-sse/utils/directResponseStartTimeout.ts`) is exactly two attempts: pooled, then fresh. When
attempt 2 — a brand-new socket that cannot possibly be a zombie pooled connection — **also** times out at 30s,
the retry logic has nothing left to try and the request fails with `DIRECT_RESPONSE_START_TIMEOUT_CODE`
(surfaced as the 504 seen in the logs). A fresh socket succeeding to *connect* but the *server* never sending a
response is exactly what "the upstream process is alive but stuck" looks like from OmniRoute's side — it can't
distinguish "GPU is wedged mid-generation" from "stale pooled socket," because both present as "nothing came
back in 30s." 46 consecutive both-attempts-failed cycles over 48 minutes is far outside what a transient stale-socket
burst (the scenario #4252/#10214 were built for) would produce; it's consistent with a sustained upstream
outage instead.
**Not currently configured in this repo**: `grep`-ing `docker-compose.yml` and `.env.example` for
`OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS` finds nothing — this deployment runs on the unmodified 30s default. (This
is also, notably, a *fourth* data point alongside the two already-documented mechanisms and the `requestQueue.maxWaitMs`/
`RATE_LIMIT_EXECUTION_TIMEOUT` timer mentioned in passing in `omniroute-account-semaphore-timeout.md` — OmniRoute
has at least four independent timeout knobs guarding different stages of a request's life, three of them
30-second-flavored by default, which is worth keeping in mind the next time an unfamiliar timeout string shows up.)
## Why the upstream itself was likely unresponsive: ROCm/legacy-rocm-build#6630
This session has **no SSH/shell access to the actual R9700 server** — there's no SSH config, and nothing in
`scripts/` does remote exec, confirmed by inspecting `scripts/update.sh` and the absence of any `~/.ssh/config`
entry for the box. So none of the following is confirmed against this specific incident's `dmesg`/`rocm-smi`/
`docker logs` output — it's the closest primary-source match to the *symptom*, not a diagnosis of *this* outage.
[ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630) ("gfx1201 R9700 ROCm 7.14
llama.cpp generation hang, MES queue failure and PSP reset -62; Vulkan passes") is an **open**, actively
investigated issue (created 2026-08-19, most recent update 2026-08-28, no fix landed) that reproduces on **the
exact same GPU this repo runs on** — Radeon AI PRO R9700, `gfx1201` — running **llama.cpp with `-fa on`** (this
repo's `llama-server` also runs `--flash-attn on`), with a controlled Vulkan-vs-ROCm A/B: Vulkan completes
normally, ROCm hangs during token generation. Direct quotes:
> "During the ROCm generation stall: GPU busy reached 100%, memory busy remained 0%, VRAM use was only about 1.1
> GB, **the container remained alive but made no output progress**."
> "when MES stops responding, **the driver can stay unaware of it indefinitely** — the failure only surfaces when
> something happens to send the next MES message... if a run is left alone after it stops making progress, the
> kernel prints nothing at all, so a hang can look like a slow workload rather than a fault."
> "the failure is probabilistic, not deterministic... a single passing run on this host does not indicate a
> healthy configuration."
The thread (12 comments as of this research pass, an AMD engineer `harkgill-amd` participating) has ruled out,
one at a time, `uni_mes=0`, `mes_log_enable=1`, ROCm 6.4.4, ROCm 7.14, ROCm 10.0.0 stable, the latest TheRock
nightly, and GFXOFF-disable — **no confirmed fix or workaround exists in the thread as of this research pass**.
A related comment on the same issue (`chrisfranson`) reports the identical MES `REMOVE_QUEUE`/MODE1-reset
signature from a **completely unrelated workload** (headless LibreOffice with OpenCL) on the same `gfx1201`
silicon, reinforcing that this is a driver/firmware-level fault under general GPU load, not something specific
to llama.cpp's request pattern.
**This is a different bug from the one already mitigated in this repo.** `docs/research/rocm-gpu-pin-and-render-group.md`
already documents and works around `ROCm/ROCm#5706` (clock/power pinned at boost whenever two concurrent HIP
contexts share the GPU — fixed via `GPU_MAX_HW_QUEUES=1`, already set on both `llama-server` and `qwen-classifier`
in `docker-compose.yml`). #5706's symptom is elevated power draw with the GPU still working; #6630's symptom is
generation fully halting with `gpu_busy=100%`/`mem_busy=0%` and MES no longer responding at all — a real hang, not
a clock-pin inefficiency. `GPU_MAX_HW_QUEUES=1` targets #5706's specific trigger (hardware-queue oversubscription
across concurrent HIP processes) and has no evidence in #6630's thread of affecting that bug — #6630 reproduces
in single-GPU, single-process benchmarks with no second HIP context involved at all, so the already-applied fix
should not be assumed to help here.
## What would actually resolve this vs. what wouldn't
- **Raising `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`** — not recommended as a fix. It would make each failed attempt
take longer before giving up (worse latency during a real hang), without addressing why the GPU stopped
responding. It's the right lever only if future evidence shows genuinely-slow-but-working responses being
mistaken for hangs (the same shape as the already-fixed `REQUEST_TIMEOUT_MS` issue in
`omniroute-non-ping-sse-stream-timeout.md`) — this incident's 46-for-46 both-attempts-failed pattern over 48
minutes doesn't fit that shape.
- **Concrete next step for whoever has server access when this recurs**: check `dmesg | grep -i amdgpu` and
`journalctl -k` on the R9700 host for `MES(...) failed to respond`, `GPU reset begin`, or `PSP resume failed`
lines matching #6630's signature, and `docker logs llama-server`/`docker logs qwen-classifier` to see whether
the process was alive-but-stuck (consistent with #6630) versus crashed/restarted (which would point elsewhere).
Capturing this during a live incident is the only way to move this from "best primary-source match" to
"confirmed root cause."
- **Monitor [ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630)** for a fix —
it's open and active (AMD engineer engaged as of 2026-08-26); no released ROCm version as of this research pass
is confirmed clean.
- **The `-UD-Q8_K_XL.gguf` classifier variant in the pasted table but absent from version control** is worth a
direct look on the server (`cat .env` / `docker inspect qwen-classifier` for the actual `LLAMA_CLASSIFIER_MODEL_FILE`
in effect) — outside this research pass's reach without server access, flagged here so it isn't lost.
## Recovery: how the hang actually clears (or doesn't) — addendum, 2026-09-15
Follow-up question: what actually recovers the socket once this hits, given it's been observed to stay wedged
for days at a time? Pulled the full comment thread on
[ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630) directly via the GitHub API
(12 comments, `angelhalo` as primary reporter, `harkgill-amd` as the responding AMD engineer, plus one
corroborating report from `chrisfranson` on unrelated hardware/workload) — the earlier research pass's source list
cited this issue but hadn't read the full thread. Three things fall directly out of it:
**There is no reliable in-band recovery.** The driver doesn't notice the hang on its own — direct quote:
"when MES stops responding, the driver can stay unaware of it indefinitely — the failure only surfaces when
something happens to send the next MES message." In a captured live hang, "every driver-managed ring is
completely idle while the GPU reports 100% busy," `dmesg` has zero amdgpu lines, and no task is in D-state —
so from the OS's perspective nothing is wrong; only sending the GPU another command (which killing/restarting
the stuck process does) triggers the driver to discover the wedge and attempt its own MODE1 reset.
**Once triggered, that reset itself is a coin flip across three documented outcomes**, not a guaranteed fix:
1. **Clean recovery** — `GPU reset succeeded, trying to resume`, PSP resumes, the card comes back (VRAM is
wiped — "VRAM is lost due to GPU reset!" — so the container needs a real restart to reload the model, a plain
process respawn isn't enough even when the reset itself works).
2. **Failed resume** — `PSP resume failed`, `GPU reset end with ret = -62` (the original report's own outcome) —
the reset attempt itself fails, leaving the GPU in a worse state than before.
3. **Full kernel soft-lockup** — `chrisfranson`'s independent report (different workload — headless LibreOffice
OpenCL, different card — RX 9070 XT, same `gfx1201` silicon) hit outcome 2 or 3 twice out of three times:
"the whole system hard-locked (kernel soft lockup pegging a CPU at ~90-100% softirq, requiring a physical
power cycle)."
`angelhalo` deliberately left one hang untouched rather than killing the process, to observe it without
contaminating the state with a reset: "I recovered only with a subsequent cold power cycle" — no reset was ever
triggered because nothing sent the GPU another message. Their standard test procedure between every single run
in this thread is "a cold power cycle (AC removed, ≥30 s)," specifically **not** a warm/soft reboot — stated
reason: "on this card a MODE1 reset takes the host down with it," meaning even the OS's own reboot path can't
be trusted to come back cleanly once this GPU is in a bad state. This is the practical answer to "why does it
stay stuck for days": nothing about the hang self-clears, `docker`'s `restart: unless-stopped` policy never
fires because the container process is alive and never exits (confirmed: this repo's `llama-server` and
`qwen-classifier` services have no `healthcheck` block at all — only `omniroute` itself does, a plain TCP
connect check on its own dashboard port, which says nothing about whether `llama-server`/`qwen-classifier` are
responding) — so a hang persists until a human notices the symptom (requests failing) and manually intervenes,
and "days" is just however long that takes to notice on a homelab box, not a property of the hang itself.
**No fix or reliable mitigation exists as of this reading (2026-08-28, the thread's latest comment).**
`harkgill-amd` (AMD) could not reproduce locally and asked for a nightly-driver retest; `angelhalo` retested and
it still failed. Every other variable tested still hangs: ROCm 6.4.4 through 10.0.0 stable, TheRock nightlies,
`amdgpu.uni_mes=0`, `cwsr_enable=0`, `mes_log_enable=1`, GFXOFF disabled, two different physical R9700 cards, both
llama.cpp and vLLM. The thread's own conclusion, as of the last comment: "a probabilistic lost-completion event"
with no known trigger to avoid and no known driver/firmware combination that's clean.
**Practical takeaway for this repo, given no upstream fix exists:**
- A restart *might* recover it, *might* make it worse (failed PSP resume), and *might* take the whole host down
requiring a physical power cycle — there's no way to know in advance which outcome a given hang will produce.
- Nothing currently watches for this automatically. Docker's `restart: unless-stopped` is the wrong tool (process
doesn't exit) — recovering automatically would need a `healthcheck` against `llama-server`'s own `/health`
endpoint (llama.cpp's built-in liveness endpoint) paired with something that acts on an `unhealthy` status,
since Docker itself doesn't restart on failed healthchecks without an external watcher (e.g. `willfarrell/autoheal`
or equivalent) — not evaluated here, flagged as a real gap, not a recommendation to implement blind: an
automated restart during a hang that's about to fail its PSP resume and lock the host could turn a
"requests are failing" incident into "the box needs a physical power cycle" automatically and unattended,
which is a real downside worth weighing against faster detection.
- Given the reset outcome is unpredictable, the safest manual recovery when this is caught live is: restart the
affected container, then immediately check `dmesg | grep -i amdgpu` for `PSP resume failed` or a soft-lockup
signature before assuming it's fixed — if either appears, a full reboot (and per this thread's own testing
practice, possibly a genuine AC power cycle rather than a warm reboot) is the next step, not a second restart
attempt.
## Caveats and open questions
- **JSON export vs. pasted table are two different, non-overlapping captures.** The JSON file is a precise 6-hour
window with full per-request detail; the pasted table is a longer, dashboard-UI-copied range with less
structure and no verifiable overlap with the JSON file's timestamps. A fresh multi-day JSON export (same
`request-logs` endpoint used to produce the file analyzed here) would let a future pass check whether the
classifier's failures and the 27B model's outage are literally simultaneous (strong evidence for a shared
GPU-level cause) or independent recurrences of the same mechanism on separate schedules.
- **Why the outage self-recovered after ~48 minutes with no observed restart is unexplained.** #6630's thread
describes hangs resolving via an explicit GPU reset (sometimes failing, requiring reboot) — not a case of a
hang clearing on its own after a fixed interval. Nothing in the available data (no server access) confirms
whether a restart happened that isn't visible from OmniRoute's logs, or whether this specific hang genuinely
self-cleared, which would be a data point *against* the #6630 hypothesis worth capturing next time.
- **Live reproduction was not attempted.** The user suggested testing tool-calls against the classifier via the
Windows-side qwen-code CLI (`C:\Users\aerli\AppData\Local\qwen-code\bin\qwen.cmd`) to try to reproduce a
"Direct response did not start" failure live. Skipped for this pass: qwen-code requires an interactive/
already-authenticated session to drive meaningfully, and deliberately trying to reproduce a GPU hang against
the shared production classifier risked a genuine 60s+ stall on infrastructure other work depends on, for
uncertain diagnostic payoff given the strength of the log-based and source-based evidence already gathered.
Worth doing deliberately, with server access on hand to capture `rocm-smi`/`dmesg` simultaneously, rather than
as a quick check from this pass.
## Sources
- `omniroute-request-logs-6h-2026-09-15.json` — OmniRoute dashboard request-log export provided by the user
(2026-09-15, 339 entries, `12:49:52Z`-`18:44:54Z`)
- User-pasted OmniRoute dashboard table (`qwen-classifier`/`qwen3-4b` failures, separate capture window)
- [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) —
`open-sse/utils/directResponseStartTimeout.ts`, `open-sse/utils/proxyFetch.ts`, `open-sse/utils/proxyDispatcher.ts`,
`tests/unit/proxyfetch-direct-response-start-timeout-10214.test.ts`, `tests/unit/proxyfetch-retry-fresh-socket-4252.test.ts`,
`open-sse/handlers/chatCore/upstreamTimeouts.ts`
- [ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630) — open R9700/gfx1201
llama.cpp generation-hang issue, full comment thread (2026-08-19 through 2026-08-28)
- [`docs/research/omniroute-account-semaphore-timeout.md`](./omniroute-account-semaphore-timeout.md) — the
first already-documented 30s OmniRoute timeout (account semaphore, 429, hardcoded)
- [`docs/research/omniroute-non-ping-sse-stream-timeout.md`](./omniroute-non-ping-sse-stream-timeout.md) — the
second already-documented timeout (first-SSE-event deadline, `REQUEST_TIMEOUT_MS`-derived)
- [`docs/research/rocm-gpu-pin-and-render-group.md`](./rocm-gpu-pin-and-render-group.md) — the already-mitigated,
*different* R9700/gfx1201 MES bug ([ROCm/ROCm#5706](https://github.com/ROCm/ROCm/issues/5706), clock-pin/power,
not a hang)
- Local `docker-compose.yml`, `.env.example` (grepped directly, confirming `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`
is unset / on the 30s default, and that `GPU_MAX_HW_QUEUES=1` is already applied to both GPU services)
- `git log --since=2026-09-14 --until=2026-09-16` (this repo, confirming zero commits during the outage window)
## Confidence / uncertainty summary
- **High confidence**: the exact mechanism and semantics of `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS` /
`directFetchWithBoundedResponseStart` (read directly from OmniRoute's own source and its two regression test
files, which spell out the intent in comments referencing the originating issues); the JSON log's chronology,
single-connection scope, and the "two ~30s attempts, both failing" shape of every 504 (computed directly from
the log file); that this timeout is unconfigured in this repo (direct grep); that no commits landed in this
repo during the outage window (direct `git log`).
- **Medium confidence**: that ROCm/legacy-rocm-build#6630 is the actual root cause of *this specific* outage.
The GPU model, driver-family symptom shape (`gpu_busy=100%`/`mem_busy=0%`, alive-but-stuck, sometimes
logged/sometimes silent), and `-fa on` usage all match closely, and it's an open/unresolved/actively-discussed
issue as of this research pass — but nothing from this incident's own `dmesg`/`rocm-smi` output was available
to confirm it directly (no SSH access), so this is the best primary-source match to the symptom, not a
confirmed diagnosis.
- **Low confidence / open**: why the outage recovered on its own after ~48 minutes with no observed restart;
whether the classifier's pasted-table failures share the exact same triggering event as the 27B outage
analyzed here (same failure signature, but no verified time-overlap between the two data sources); the
provenance of the `-UD-Q8_K_XL.gguf` classifier variant seen in the pasted table but absent from version
control.
## Live test, 2026-09-15: the classifier is measurably too slow at its own documented worst case — independent of any hang
Before scoping a fix, tested the live `qwen-classifier` backend directly against realistic worst-case load,
per the user's request to gather fresh evidence rather than design blind. Two attempts to reproduce this through
qwen-code itself first surfaced an unrelated, separately-useful finding; the direct backend test below is what
actually answered the question.
### qwen-code's own headless mode never reaches the classifier
Ran `qwen --approval-mode auto <prompt>` (positional/one-shot, non-interactive) from the Windows-side install
(`C:\Users\aerli\AppData\Local\qwen-code\bin\qwen.cmd`, which has both `fastModel` and the `omniroute-search` MCP
server already configured), asking it to run a shell `dir` and use the web-search MCP tool. Both attempts hit a
wall before any classifier request was even sent:
- The MCP tool call was refused outright: `Warning: Tool "mcp__omniroute-search__search" requires user approval
but cannot execute in non-interactive mode. ... use the -y flag (YOLO mode)`.
- The shell tool: the model itself reported `run_shell_command` as "not registered" in this session and silently
substituted a read-only `glob` call instead — no approval prompt, no classifier call, no system warning printed
(unlike the MCP case), across two separate clean runs.
**Conclusion: one-shot headless `qwen <prompt>` invocations don't exercise Auto Mode's classifier at all for
approval-requiring tools** — they're declined or silently rerouted before the classifier ever gets a request.
The classifier only fires in a genuinely interactive session, where it substitutes for the human's live approval
decision. This wasn't previously documented anywhere in this repo and is worth keeping in mind: headless qwen-code
testing is not a valid way to probe classifier behavior, live or otherwise. (Not investigated further: whether
`qwen serve`/`--input-format stream-json` headless-agent modes behave differently — plausible, since they're
built for exactly this kind of automation, but out of scope for this pass.)
### Direct backend test: real classifier latency at realistic token counts
Given headless qwen-code couldn't drive this, sent shaped classifier requests straight to
`http://proxy-ai.home/v1/chat/completions` (model `qwen3-4b//models/Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`,
confirmed present in `GET /v1/models`) using a temporary scoped API key, mimicking the `{shouldBlock}`
JSON-verdict shape and the `MAX_TRANSCRIPT_MESSAGES=40` / `MAX_HISTORICAL_ACTION_CHARS=4000` structure this
repo's own `fast-model-choice.md` already read out of qwen-code's `classifier-transcript.ts` source. Five calls,
in order:
| # | Prompt tokens | `cached_tokens` | Wall time | Notes |
|---|---|---|---|---|
| 1 | 1,000 | 3 | 2.24s | Small prompt, genuinely fresh — healthy baseline. |
| 2 | 51,234 | 51,233 | **64.21s** | First send of a large synthetic worst-case transcript (~40 highly self-similar 4K-char "historical action" blocks). Near-total cache hit reported, yet still the slowest call — see caveat below. |
| 3 | 51,234 | 51,233 | 0.07s | **Identical repeat of #2.** Same `chatcmpl-...` id as #2 came back — this is OmniRoute short-circuiting an exact-duplicate request via a proxy-level response cache, not fresh inference. Confirms #2/#3's `cached_tokens` field is not a reliable proxy for wall-clock latency on its own. |
| 4 | 1,000 | 3 | 0.03s | Identical repeat of #1 — same id, same response-cache short-circuit. |
| 5 | 30,958 | 919 | **28.66s** | Fresh, non-repeated worst-case-shaped transcript (different random content, no internal self-similarity to trigger cache effects). Mostly-uncached (919/30,958) — this is the clean data point. |
Call #5 is the one to trust: **~31K genuinely-fresh prompt tokens took 28.7 seconds** on the current
`--n-gpu-layers 28` (of 36) / `--cache-type-k/v q4_0` / `--parallel 1` configuration, with the backend otherwise
idle and healthy (no hang in progress). Extrapolating that rate to the repo's own documented worst case (a real
15,116-token call observed live per `fast-model-choice.md`, and a theoretical ceiling around 40-50K tokens per
`classifier-transcript.ts`'s limits) puts a genuine worst-case classifier call at **roughly 30-65 seconds of
normal, non-hung processing time** — consistent with call #2's 64.21s, even though that call's own cache
metadata is too muddied by internal prompt self-similarity to use as a second clean sample.
**This directly overlaps both binding timeouts**: qwen-code's own client-side classifier stage timeout
(`stage1Ms`/`stage2Ms`, `60000` each in the Windows-side `settings.json` observed this session, `30000`/`60000`
in the WSL-side one) and `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`'s 30-second-per-attempt window documented above.
A worst-case classifier call landing in the 30-65s range will **routinely** trip one or both of these timeouts
on its own, with the backend never having hung at all — the exact same 499/504/"Request aborted" shape as the
GPU-hang hypothesis produces, but from an entirely mundane, deterministic cause: **the classifier's current
configuration is simply too slow for the request sizes qwen-code's own classifier-transcript design allows.**
### What this changes
This doesn't rule out ROCm/legacy-rocm-build#6630 — the 27B model's 48-minute total outage (zero successes, not
just slow ones) doesn't fit a "just slow" explanation, and remains best matched by a real GPU hang. But it does
mean **the classifier's own recurring failures (the pasted-table evidence) very plausibly have a second,
independent, non-probabilistic cause that a healthcheck/restart-on-hang design wouldn't fix at all** — restarting
a classifier that's merely slow-but-working at worst-case load just interrupts a call that would have succeeded,
and would fire repeatedly under normal peak usage, not just during a rare hang. Any fix that only targets "detect
and recover from an unresponsive GPU" leaves this second failure mode untouched. Two independent levers worth
weighing before finalizing a scope: raising the classifier's own timeouts to match its real worst-case latency
(cheap, immediate, but does nothing for actual hangs), and/or speeding up the classifier itself (full GPU offload
if VRAM allows, a faster quant, or capping the transcript size client-side) to bring worst-case latency back
under the existing timeouts.
**Not investigated in this pass**: whether call #2's 64.21s (vs. call #5's extrapolated ~45-48s at a similar
token count) reflects genuine non-linear slowdown at the very largest context sizes, real concurrent contention
from other production traffic sharing the same `--parallel 1` slot during the test, or is just noise from a
single sample each — worth a few more clean, uniquely-content, worst-case-sized calls at different times of day
before treating either number as precise.