From ed90256f437a96eb858a3177e45aaea0ba1253c6 Mon Sep 17 00:00:00 2001 From: ArthurErlich Date: Tue, 15 Sep 2026 23:31:58 +0200 Subject: [PATCH] docs(research): trace 48-min outage to third OmniRoute timeout and open ROCm#6630 GPU hang --- ...rect-response-timeout-outage-2026-09-15.md | 397 ++++++++++++++++++ 1 file changed, 397 insertions(+) create mode 100644 docs/research/omniroute-direct-response-timeout-outage-2026-09-15.md diff --git a/docs/research/omniroute-direct-response-timeout-outage-2026-09-15.md b/docs/research/omniroute-direct-response-timeout-outage-2026-09-15.md new file mode 100644 index 0000000..302e439 --- /dev/null +++ b/docs/research/omniroute-direct-response-timeout-outage-2026-09-15.md @@ -0,0 +1,397 @@ +# A 48-minute total outage on `qwen3.8-27b-local`, and the third OmniRoute timeout mechanism this repo hadn't documented yet + +**Date:** 2026-09-15 + +**Verdict:** The error — `"[504]: Direct response did not start within 30000ms — retrying on a fresh socket"` — +comes from a **third, previously-undocumented OmniRoute timeout mechanism** (`OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`, +default 30s), distinct from both timeouts already recorded in +[`omniroute-account-semaphore-timeout.md`](./omniroute-account-semaphore-timeout.md) and +[`omniroute-non-ping-sse-stream-timeout.md`](./omniroute-non-ping-sse-stream-timeout.md). It exists specifically to +recover from a *stale pooled TCP socket* by retrying once on a brand-new connection — but in the incident analyzed +here, **both** the original attempt and the fresh-socket retry timed out, repeatedly, for 46 requests over 48 +straight minutes with zero successes. That pattern rules out a stale-socket explanation (a fresh socket bypasses +the pool entirely) and points instead at the upstream itself — `llama-server`, or the R9700 GPU underneath it — +being genuinely unresponsive for the whole window. The best primary-source match for that symptom is an **open, +still-unresolved AMD ROCm bug specific to this exact GPU** ([ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630)): +an MES-firmware hang during generation on `gfx1201`/R9700 that leaves the process alive but stuck, sometimes for +no logged reason at all. No config change fixes this — raising the 30s timeout only makes each failed attempt +take longer to give up, it doesn't un-wedge a hung GPU. + +## The evidence + +**Source:** `omniroute-request-logs-6h-2026-09-15.json`, a 339-entry OmniRoute request-log export the user pulled +from the dashboard, covering `2026-09-15T12:49:52Z`–`18:44:54Z`. Every entry with a non-200 status (66 of them) +is a `POST /v1/chat/completions` against `qwen3.8-27b-local` (`/models/Qwen3.8-27B-UD-Q4_K_XL.gguf`), all on the +same `connectionId` (`649a2d3e-7527-488e-9b8a-dc4ac2624176`) and the same `provider` +(`openai-compatible-chat-a7bda643-6687-41f4-b75d-fd2cab746874`) — a single upstream connection, not a fan-out +artifact. + +Sorting every error by timestamp shows **one continuous outage**, not scattered slow requests: + +- Last successful `/v1/chat/completions` before the outage: `15:18:12.285Z` +- **First failure:** `15:25:06.415Z` — status 504, `"[504]: Direct response did not start within 30000ms — + retrying on a fresh socket"`, duration `60186ms` +- **Every single `/v1/chat/completions` attempt** from `15:25:06.415Z` through `16:13:29.502Z` failed — 46× 504 + (all clustered `60021`-`60324ms`, i.e. two back-to-back 30s attempts, both failing) interleaved with 20× 499 + (`"Request aborted"` / `"Client disconnected: request_signal_aborted"`, durations `1.4s`-`99.97s` — these are + qwen-code giving up client-side while OmniRoute was still mid-retry) +- **First successful recovery:** `16:13:50.388Z`, `20873ms` — 21 minutes after the last failure attempt cluster, + i.e. the very next attempt after the outage window succeeded normally +- No successful `/v1/chat/completions` call appears anywhere inside the `15:25:06Z`-`16:13:29Z` window — confirmed + by filtering all 154 `/v1/chat/completions` log entries in that range: every one is 504 or 499. + +`git log --since=2026-09-14 --until=2026-09-16` shows **zero commits** in this repo on 2026-09-15 — the outage +correlates with no deploy, `scripts/update.sh` run, or config change on this end. + +**A second, independent data point** (see caveat below): the user also pasted a large raw text table, copied +directly from OmniRoute's dashboard UI rather than the JSON export, showing the **`qwen-classifier`** connection +(`qwen3-4b`, both the `-UD-Q4_K_XL.gguf` file this repo's `docker-compose.yml` currently defaults to, and a +`-UD-Q8_K_XL.gguf` variant that appears **nowhere** in this repo's checked-in `docker-compose.yml`/`.env.example` +— either tested by hand against `LLAMA_CLASSIFIER_MODEL_FILE` outside version control, or evidence of drift worth +checking directly on the server) failing repeatedly across two accounts (`Haylan`, `qwen-cli-main`), with the +same signature: `TI: 0|TO: 0` (zero tokens either direction — failed before generating anything) and durations +clustering at exactly `60.0`-`60.4s` for the 504s. That's the identical two-attempts-at-30s-each shape as the 27B +outage above, strongly suggesting the same underlying mechanism, though — important caveat — **this table is not +present in the 6-hour JSON export and covers a different, longer time range** (timestamps back to `01:21` and +`23:54` on unspecified dates), so it cannot be directly time-correlated against the 27B outage above. Treat it as +corroborating evidence that this failure mode recurs on both local model connections, not as proof they failed at +the same moment. + +## Root cause: `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`, a mechanism built for a different problem + +Confirmed directly against OmniRoute's own source, [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) +(same repo the two existing timeout docs already cite) — `open-sse/utils/directResponseStartTimeout.ts`: + +```ts +const DEFAULT_DIRECT_HEADERS_TIMEOUT_MS = 30_000; +const DIRECT_RESPONSE_START_TIMEOUT_CODE = "DIRECT_RESPONSE_START_TIMEOUT"; + +export function resolveDirectHeadersTimeoutMs( + env: Record = process.env +): number { + const raw = env.OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS; + if (raw == null || raw.trim() === "") return DEFAULT_DIRECT_HEADERS_TIMEOUT_MS; + ... +} + +function createDirectResponseStartTimeout(timeoutMs: number): Error & { code: string } { + const err = new Error( + `Direct response did not start within ${timeoutMs}ms — retrying on a fresh socket` + ) as Error & { code: string }; + ... +} +``` + +"Direct" here means **direct (no-proxy) egress** — confirmed in `open-sse/utils/proxyFetch.ts`, which routes any +connection with no configured upstream HTTP proxy through this path (as opposed to OmniRoute's separate +proxy/relay egress paths). Every local connection in this repo (`llama-server`, `qwen-classifier`, both reached +over the `ai-stack` Docker network with no proxy) is "direct" — so this timeout mechanism governs **every** +request to either local model, streaming or non-streaming alike, not just non-streaming JSON responses as the +name might suggest. + +### Why it exists (and why it didn't help here) + +Two OmniRoute issues, both with dedicated regression tests in the repo, explain the actual design intent: + +- **#4252** (`tests/unit/proxyfetch-retry-fresh-socket-4252.test.ts`): "Undici dispatcher fails on direct provider + requests in 502 bursts" — the default direct dispatcher pools keep-alive sockets; some upstreams silently close + idle pooled sockets, so the next request reusing one fails with `UND_ERR_SOCKET`. Fix: retry once on a **fresh, + no-keep-alive dispatcher** (`getRetryDispatcher()`, a different instance from `getDefaultDispatcher()`) so the + retry can't grab another already-dead pooled socket. +- **#10214** (`tests/unit/proxyfetch-direct-response-start-timeout-10214.test.ts`): "Direct (no-proxy) requests + stall on a silently-dropped pooled keep-alive socket until the caller's deadline or a service restart" — the + harder case: a pooled socket that dies **without even an error**, just silence. Undici's `headersTimeout` + default (600s) is far too slow to catch this in practice, and the existing #4252 retry never fires because no + error is thrown to trigger it. The fix bounds each direct attempt's response-start wait to + `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS` (30s default) via `directFetchWithBoundedResponseStart`, and retries + **once** on the same fresh no-keep-alive dispatcher from #4252 when that bound is hit. + +Both fixes assume the *socket* is the problem, not the upstream. `directFetchWithBoundedResponseStart`'s own +implementation (`open-sse/utils/directResponseStartTimeout.ts`) is exactly two attempts: pooled, then fresh. When +attempt 2 — a brand-new socket that cannot possibly be a zombie pooled connection — **also** times out at 30s, +the retry logic has nothing left to try and the request fails with `DIRECT_RESPONSE_START_TIMEOUT_CODE` +(surfaced as the 504 seen in the logs). A fresh socket succeeding to *connect* but the *server* never sending a +response is exactly what "the upstream process is alive but stuck" looks like from OmniRoute's side — it can't +distinguish "GPU is wedged mid-generation" from "stale pooled socket," because both present as "nothing came +back in 30s." 46 consecutive both-attempts-failed cycles over 48 minutes is far outside what a transient stale-socket +burst (the scenario #4252/#10214 were built for) would produce; it's consistent with a sustained upstream +outage instead. + +**Not currently configured in this repo**: `grep`-ing `docker-compose.yml` and `.env.example` for +`OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS` finds nothing — this deployment runs on the unmodified 30s default. (This +is also, notably, a *fourth* data point alongside the two already-documented mechanisms and the `requestQueue.maxWaitMs`/ +`RATE_LIMIT_EXECUTION_TIMEOUT` timer mentioned in passing in `omniroute-account-semaphore-timeout.md` — OmniRoute +has at least four independent timeout knobs guarding different stages of a request's life, three of them +30-second-flavored by default, which is worth keeping in mind the next time an unfamiliar timeout string shows up.) + +## Why the upstream itself was likely unresponsive: ROCm/legacy-rocm-build#6630 + +This session has **no SSH/shell access to the actual R9700 server** — there's no SSH config, and nothing in +`scripts/` does remote exec, confirmed by inspecting `scripts/update.sh` and the absence of any `~/.ssh/config` +entry for the box. So none of the following is confirmed against this specific incident's `dmesg`/`rocm-smi`/ +`docker logs` output — it's the closest primary-source match to the *symptom*, not a diagnosis of *this* outage. + +[ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630) ("gfx1201 R9700 ROCm 7.14 +llama.cpp generation hang, MES queue failure and PSP reset -62; Vulkan passes") is an **open**, actively +investigated issue (created 2026-08-19, most recent update 2026-08-28, no fix landed) that reproduces on **the +exact same GPU this repo runs on** — Radeon AI PRO R9700, `gfx1201` — running **llama.cpp with `-fa on`** (this +repo's `llama-server` also runs `--flash-attn on`), with a controlled Vulkan-vs-ROCm A/B: Vulkan completes +normally, ROCm hangs during token generation. Direct quotes: + +> "During the ROCm generation stall: GPU busy reached 100%, memory busy remained 0%, VRAM use was only about 1.1 +> GB, **the container remained alive but made no output progress**." + +> "when MES stops responding, **the driver can stay unaware of it indefinitely** — the failure only surfaces when +> something happens to send the next MES message... if a run is left alone after it stops making progress, the +> kernel prints nothing at all, so a hang can look like a slow workload rather than a fault." + +> "the failure is probabilistic, not deterministic... a single passing run on this host does not indicate a +> healthy configuration." + +The thread (12 comments as of this research pass, an AMD engineer `harkgill-amd` participating) has ruled out, +one at a time, `uni_mes=0`, `mes_log_enable=1`, ROCm 6.4.4, ROCm 7.14, ROCm 10.0.0 stable, the latest TheRock +nightly, and GFXOFF-disable — **no confirmed fix or workaround exists in the thread as of this research pass**. +A related comment on the same issue (`chrisfranson`) reports the identical MES `REMOVE_QUEUE`/MODE1-reset +signature from a **completely unrelated workload** (headless LibreOffice with OpenCL) on the same `gfx1201` +silicon, reinforcing that this is a driver/firmware-level fault under general GPU load, not something specific +to llama.cpp's request pattern. + +**This is a different bug from the one already mitigated in this repo.** `docs/research/rocm-gpu-pin-and-render-group.md` +already documents and works around `ROCm/ROCm#5706` (clock/power pinned at boost whenever two concurrent HIP +contexts share the GPU — fixed via `GPU_MAX_HW_QUEUES=1`, already set on both `llama-server` and `qwen-classifier` +in `docker-compose.yml`). #5706's symptom is elevated power draw with the GPU still working; #6630's symptom is +generation fully halting with `gpu_busy=100%`/`mem_busy=0%` and MES no longer responding at all — a real hang, not +a clock-pin inefficiency. `GPU_MAX_HW_QUEUES=1` targets #5706's specific trigger (hardware-queue oversubscription +across concurrent HIP processes) and has no evidence in #6630's thread of affecting that bug — #6630 reproduces +in single-GPU, single-process benchmarks with no second HIP context involved at all, so the already-applied fix +should not be assumed to help here. + +## What would actually resolve this vs. what wouldn't + +- **Raising `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`** — not recommended as a fix. It would make each failed attempt + take longer before giving up (worse latency during a real hang), without addressing why the GPU stopped + responding. It's the right lever only if future evidence shows genuinely-slow-but-working responses being + mistaken for hangs (the same shape as the already-fixed `REQUEST_TIMEOUT_MS` issue in + `omniroute-non-ping-sse-stream-timeout.md`) — this incident's 46-for-46 both-attempts-failed pattern over 48 + minutes doesn't fit that shape. +- **Concrete next step for whoever has server access when this recurs**: check `dmesg | grep -i amdgpu` and + `journalctl -k` on the R9700 host for `MES(...) failed to respond`, `GPU reset begin`, or `PSP resume failed` + lines matching #6630's signature, and `docker logs llama-server`/`docker logs qwen-classifier` to see whether + the process was alive-but-stuck (consistent with #6630) versus crashed/restarted (which would point elsewhere). + Capturing this during a live incident is the only way to move this from "best primary-source match" to + "confirmed root cause." +- **Monitor [ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630)** for a fix — + it's open and active (AMD engineer engaged as of 2026-08-26); no released ROCm version as of this research pass + is confirmed clean. +- **The `-UD-Q8_K_XL.gguf` classifier variant in the pasted table but absent from version control** is worth a + direct look on the server (`cat .env` / `docker inspect qwen-classifier` for the actual `LLAMA_CLASSIFIER_MODEL_FILE` + in effect) — outside this research pass's reach without server access, flagged here so it isn't lost. + +## Recovery: how the hang actually clears (or doesn't) — addendum, 2026-09-15 + +Follow-up question: what actually recovers the socket once this hits, given it's been observed to stay wedged +for days at a time? Pulled the full comment thread on +[ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630) directly via the GitHub API +(12 comments, `angelhalo` as primary reporter, `harkgill-amd` as the responding AMD engineer, plus one +corroborating report from `chrisfranson` on unrelated hardware/workload) — the earlier research pass's source list +cited this issue but hadn't read the full thread. Three things fall directly out of it: + +**There is no reliable in-band recovery.** The driver doesn't notice the hang on its own — direct quote: +"when MES stops responding, the driver can stay unaware of it indefinitely — the failure only surfaces when +something happens to send the next MES message." In a captured live hang, "every driver-managed ring is +completely idle while the GPU reports 100% busy," `dmesg` has zero amdgpu lines, and no task is in D-state — +so from the OS's perspective nothing is wrong; only sending the GPU another command (which killing/restarting +the stuck process does) triggers the driver to discover the wedge and attempt its own MODE1 reset. + +**Once triggered, that reset itself is a coin flip across three documented outcomes**, not a guaranteed fix: +1. **Clean recovery** — `GPU reset succeeded, trying to resume`, PSP resumes, the card comes back (VRAM is + wiped — "VRAM is lost due to GPU reset!" — so the container needs a real restart to reload the model, a plain + process respawn isn't enough even when the reset itself works). +2. **Failed resume** — `PSP resume failed`, `GPU reset end with ret = -62` (the original report's own outcome) — + the reset attempt itself fails, leaving the GPU in a worse state than before. +3. **Full kernel soft-lockup** — `chrisfranson`'s independent report (different workload — headless LibreOffice + OpenCL, different card — RX 9070 XT, same `gfx1201` silicon) hit outcome 2 or 3 twice out of three times: + "the whole system hard-locked (kernel soft lockup pegging a CPU at ~90-100% softirq, requiring a physical + power cycle)." + +`angelhalo` deliberately left one hang untouched rather than killing the process, to observe it without +contaminating the state with a reset: "I recovered only with a subsequent cold power cycle" — no reset was ever +triggered because nothing sent the GPU another message. Their standard test procedure between every single run +in this thread is "a cold power cycle (AC removed, ≥30 s)," specifically **not** a warm/soft reboot — stated +reason: "on this card a MODE1 reset takes the host down with it," meaning even the OS's own reboot path can't +be trusted to come back cleanly once this GPU is in a bad state. This is the practical answer to "why does it +stay stuck for days": nothing about the hang self-clears, `docker`'s `restart: unless-stopped` policy never +fires because the container process is alive and never exits (confirmed: this repo's `llama-server` and +`qwen-classifier` services have no `healthcheck` block at all — only `omniroute` itself does, a plain TCP +connect check on its own dashboard port, which says nothing about whether `llama-server`/`qwen-classifier` are +responding) — so a hang persists until a human notices the symptom (requests failing) and manually intervenes, +and "days" is just however long that takes to notice on a homelab box, not a property of the hang itself. + +**No fix or reliable mitigation exists as of this reading (2026-08-28, the thread's latest comment).** +`harkgill-amd` (AMD) could not reproduce locally and asked for a nightly-driver retest; `angelhalo` retested and +it still failed. Every other variable tested still hangs: ROCm 6.4.4 through 10.0.0 stable, TheRock nightlies, +`amdgpu.uni_mes=0`, `cwsr_enable=0`, `mes_log_enable=1`, GFXOFF disabled, two different physical R9700 cards, both +llama.cpp and vLLM. The thread's own conclusion, as of the last comment: "a probabilistic lost-completion event" +with no known trigger to avoid and no known driver/firmware combination that's clean. + +**Practical takeaway for this repo, given no upstream fix exists:** +- A restart *might* recover it, *might* make it worse (failed PSP resume), and *might* take the whole host down + requiring a physical power cycle — there's no way to know in advance which outcome a given hang will produce. +- Nothing currently watches for this automatically. Docker's `restart: unless-stopped` is the wrong tool (process + doesn't exit) — recovering automatically would need a `healthcheck` against `llama-server`'s own `/health` + endpoint (llama.cpp's built-in liveness endpoint) paired with something that acts on an `unhealthy` status, + since Docker itself doesn't restart on failed healthchecks without an external watcher (e.g. `willfarrell/autoheal` + or equivalent) — not evaluated here, flagged as a real gap, not a recommendation to implement blind: an + automated restart during a hang that's about to fail its PSP resume and lock the host could turn a + "requests are failing" incident into "the box needs a physical power cycle" automatically and unattended, + which is a real downside worth weighing against faster detection. +- Given the reset outcome is unpredictable, the safest manual recovery when this is caught live is: restart the + affected container, then immediately check `dmesg | grep -i amdgpu` for `PSP resume failed` or a soft-lockup + signature before assuming it's fixed — if either appears, a full reboot (and per this thread's own testing + practice, possibly a genuine AC power cycle rather than a warm reboot) is the next step, not a second restart + attempt. + +## Caveats and open questions + +- **JSON export vs. pasted table are two different, non-overlapping captures.** The JSON file is a precise 6-hour + window with full per-request detail; the pasted table is a longer, dashboard-UI-copied range with less + structure and no verifiable overlap with the JSON file's timestamps. A fresh multi-day JSON export (same + `request-logs` endpoint used to produce the file analyzed here) would let a future pass check whether the + classifier's failures and the 27B model's outage are literally simultaneous (strong evidence for a shared + GPU-level cause) or independent recurrences of the same mechanism on separate schedules. +- **Why the outage self-recovered after ~48 minutes with no observed restart is unexplained.** #6630's thread + describes hangs resolving via an explicit GPU reset (sometimes failing, requiring reboot) — not a case of a + hang clearing on its own after a fixed interval. Nothing in the available data (no server access) confirms + whether a restart happened that isn't visible from OmniRoute's logs, or whether this specific hang genuinely + self-cleared, which would be a data point *against* the #6630 hypothesis worth capturing next time. +- **Live reproduction was not attempted.** The user suggested testing tool-calls against the classifier via the + Windows-side qwen-code CLI (`C:\Users\aerli\AppData\Local\qwen-code\bin\qwen.cmd`) to try to reproduce a + "Direct response did not start" failure live. Skipped for this pass: qwen-code requires an interactive/ + already-authenticated session to drive meaningfully, and deliberately trying to reproduce a GPU hang against + the shared production classifier risked a genuine 60s+ stall on infrastructure other work depends on, for + uncertain diagnostic payoff given the strength of the log-based and source-based evidence already gathered. + Worth doing deliberately, with server access on hand to capture `rocm-smi`/`dmesg` simultaneously, rather than + as a quick check from this pass. + +## Sources + +- `omniroute-request-logs-6h-2026-09-15.json` — OmniRoute dashboard request-log export provided by the user + (2026-09-15, 339 entries, `12:49:52Z`-`18:44:54Z`) +- User-pasted OmniRoute dashboard table (`qwen-classifier`/`qwen3-4b` failures, separate capture window) +- [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) — + `open-sse/utils/directResponseStartTimeout.ts`, `open-sse/utils/proxyFetch.ts`, `open-sse/utils/proxyDispatcher.ts`, + `tests/unit/proxyfetch-direct-response-start-timeout-10214.test.ts`, `tests/unit/proxyfetch-retry-fresh-socket-4252.test.ts`, + `open-sse/handlers/chatCore/upstreamTimeouts.ts` +- [ROCm/legacy-rocm-build#6630](https://github.com/ROCm/legacy-rocm-build/issues/6630) — open R9700/gfx1201 + llama.cpp generation-hang issue, full comment thread (2026-08-19 through 2026-08-28) +- [`docs/research/omniroute-account-semaphore-timeout.md`](./omniroute-account-semaphore-timeout.md) — the + first already-documented 30s OmniRoute timeout (account semaphore, 429, hardcoded) +- [`docs/research/omniroute-non-ping-sse-stream-timeout.md`](./omniroute-non-ping-sse-stream-timeout.md) — the + second already-documented timeout (first-SSE-event deadline, `REQUEST_TIMEOUT_MS`-derived) +- [`docs/research/rocm-gpu-pin-and-render-group.md`](./rocm-gpu-pin-and-render-group.md) — the already-mitigated, + *different* R9700/gfx1201 MES bug ([ROCm/ROCm#5706](https://github.com/ROCm/ROCm/issues/5706), clock-pin/power, + not a hang) +- Local `docker-compose.yml`, `.env.example` (grepped directly, confirming `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS` + is unset / on the 30s default, and that `GPU_MAX_HW_QUEUES=1` is already applied to both GPU services) +- `git log --since=2026-09-14 --until=2026-09-16` (this repo, confirming zero commits during the outage window) + +## Confidence / uncertainty summary + +- **High confidence**: the exact mechanism and semantics of `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS` / + `directFetchWithBoundedResponseStart` (read directly from OmniRoute's own source and its two regression test + files, which spell out the intent in comments referencing the originating issues); the JSON log's chronology, + single-connection scope, and the "two ~30s attempts, both failing" shape of every 504 (computed directly from + the log file); that this timeout is unconfigured in this repo (direct grep); that no commits landed in this + repo during the outage window (direct `git log`). +- **Medium confidence**: that ROCm/legacy-rocm-build#6630 is the actual root cause of *this specific* outage. + The GPU model, driver-family symptom shape (`gpu_busy=100%`/`mem_busy=0%`, alive-but-stuck, sometimes + logged/sometimes silent), and `-fa on` usage all match closely, and it's an open/unresolved/actively-discussed + issue as of this research pass — but nothing from this incident's own `dmesg`/`rocm-smi` output was available + to confirm it directly (no SSH access), so this is the best primary-source match to the symptom, not a + confirmed diagnosis. +- **Low confidence / open**: why the outage recovered on its own after ~48 minutes with no observed restart; + whether the classifier's pasted-table failures share the exact same triggering event as the 27B outage + analyzed here (same failure signature, but no verified time-overlap between the two data sources); the + provenance of the `-UD-Q8_K_XL.gguf` classifier variant seen in the pasted table but absent from version + control. + +## Live test, 2026-09-15: the classifier is measurably too slow at its own documented worst case — independent of any hang + +Before scoping a fix, tested the live `qwen-classifier` backend directly against realistic worst-case load, +per the user's request to gather fresh evidence rather than design blind. Two attempts to reproduce this through +qwen-code itself first surfaced an unrelated, separately-useful finding; the direct backend test below is what +actually answered the question. + +### qwen-code's own headless mode never reaches the classifier + +Ran `qwen --approval-mode auto ` (positional/one-shot, non-interactive) from the Windows-side install +(`C:\Users\aerli\AppData\Local\qwen-code\bin\qwen.cmd`, which has both `fastModel` and the `omniroute-search` MCP +server already configured), asking it to run a shell `dir` and use the web-search MCP tool. Both attempts hit a +wall before any classifier request was even sent: + +- The MCP tool call was refused outright: `Warning: Tool "mcp__omniroute-search__search" requires user approval + but cannot execute in non-interactive mode. ... use the -y flag (YOLO mode)`. +- The shell tool: the model itself reported `run_shell_command` as "not registered" in this session and silently + substituted a read-only `glob` call instead — no approval prompt, no classifier call, no system warning printed + (unlike the MCP case), across two separate clean runs. + +**Conclusion: one-shot headless `qwen ` invocations don't exercise Auto Mode's classifier at all for +approval-requiring tools** — they're declined or silently rerouted before the classifier ever gets a request. +The classifier only fires in a genuinely interactive session, where it substitutes for the human's live approval +decision. This wasn't previously documented anywhere in this repo and is worth keeping in mind: headless qwen-code +testing is not a valid way to probe classifier behavior, live or otherwise. (Not investigated further: whether +`qwen serve`/`--input-format stream-json` headless-agent modes behave differently — plausible, since they're +built for exactly this kind of automation, but out of scope for this pass.) + +### Direct backend test: real classifier latency at realistic token counts + +Given headless qwen-code couldn't drive this, sent shaped classifier requests straight to +`http://proxy-ai.home/v1/chat/completions` (model `qwen3-4b//models/Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf`, +confirmed present in `GET /v1/models`) using a temporary scoped API key, mimicking the `{shouldBlock}` +JSON-verdict shape and the `MAX_TRANSCRIPT_MESSAGES=40` / `MAX_HISTORICAL_ACTION_CHARS=4000` structure this +repo's own `fast-model-choice.md` already read out of qwen-code's `classifier-transcript.ts` source. Five calls, +in order: + +| # | Prompt tokens | `cached_tokens` | Wall time | Notes | +|---|---|---|---|---| +| 1 | 1,000 | 3 | 2.24s | Small prompt, genuinely fresh — healthy baseline. | +| 2 | 51,234 | 51,233 | **64.21s** | First send of a large synthetic worst-case transcript (~40 highly self-similar 4K-char "historical action" blocks). Near-total cache hit reported, yet still the slowest call — see caveat below. | +| 3 | 51,234 | 51,233 | 0.07s | **Identical repeat of #2.** Same `chatcmpl-...` id as #2 came back — this is OmniRoute short-circuiting an exact-duplicate request via a proxy-level response cache, not fresh inference. Confirms #2/#3's `cached_tokens` field is not a reliable proxy for wall-clock latency on its own. | +| 4 | 1,000 | 3 | 0.03s | Identical repeat of #1 — same id, same response-cache short-circuit. | +| 5 | 30,958 | 919 | **28.66s** | Fresh, non-repeated worst-case-shaped transcript (different random content, no internal self-similarity to trigger cache effects). Mostly-uncached (919/30,958) — this is the clean data point. | + +Call #5 is the one to trust: **~31K genuinely-fresh prompt tokens took 28.7 seconds** on the current +`--n-gpu-layers 28` (of 36) / `--cache-type-k/v q4_0` / `--parallel 1` configuration, with the backend otherwise +idle and healthy (no hang in progress). Extrapolating that rate to the repo's own documented worst case (a real +15,116-token call observed live per `fast-model-choice.md`, and a theoretical ceiling around 40-50K tokens per +`classifier-transcript.ts`'s limits) puts a genuine worst-case classifier call at **roughly 30-65 seconds of +normal, non-hung processing time** — consistent with call #2's 64.21s, even though that call's own cache +metadata is too muddied by internal prompt self-similarity to use as a second clean sample. + +**This directly overlaps both binding timeouts**: qwen-code's own client-side classifier stage timeout +(`stage1Ms`/`stage2Ms`, `60000` each in the Windows-side `settings.json` observed this session, `30000`/`60000` +in the WSL-side one) and `OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS`'s 30-second-per-attempt window documented above. +A worst-case classifier call landing in the 30-65s range will **routinely** trip one or both of these timeouts +on its own, with the backend never having hung at all — the exact same 499/504/"Request aborted" shape as the +GPU-hang hypothesis produces, but from an entirely mundane, deterministic cause: **the classifier's current +configuration is simply too slow for the request sizes qwen-code's own classifier-transcript design allows.** + +### What this changes + +This doesn't rule out ROCm/legacy-rocm-build#6630 — the 27B model's 48-minute total outage (zero successes, not +just slow ones) doesn't fit a "just slow" explanation, and remains best matched by a real GPU hang. But it does +mean **the classifier's own recurring failures (the pasted-table evidence) very plausibly have a second, +independent, non-probabilistic cause that a healthcheck/restart-on-hang design wouldn't fix at all** — restarting +a classifier that's merely slow-but-working at worst-case load just interrupts a call that would have succeeded, +and would fire repeatedly under normal peak usage, not just during a rare hang. Any fix that only targets "detect +and recover from an unresponsive GPU" leaves this second failure mode untouched. Two independent levers worth +weighing before finalizing a scope: raising the classifier's own timeouts to match its real worst-case latency +(cheap, immediate, but does nothing for actual hangs), and/or speeding up the classifier itself (full GPU offload +if VRAM allows, a faster quant, or capping the transcript size client-side) to bring worst-case latency back +under the existing timeouts. + +**Not investigated in this pass**: whether call #2's 64.21s (vs. call #5's extrapolated ~45-48s at a similar +token count) reflects genuine non-linear slowdown at the very largest context sizes, real concurrent contention +from other production traffic sharing the same `--parallel 1` slot during the test, or is just noise from a +single sample each — worth a few more clean, uniquely-content, worst-case-sized calls at different times of day +before treating either number as precise.