Files
LLM-Server/docs/research/omniroute-direct-response-timeout-outage-2026-09-15.md

32 KiB
Raw Permalink Blame History

A 48-minute total outage on qwen3.8-27b-local, and the third OmniRoute timeout mechanism this repo hadn't documented yet

Date: 2026-09-15

Verdict: The error — "[504]: Direct response did not start within 30000ms — retrying on a fresh socket" — comes from a third, previously-undocumented OmniRoute timeout mechanism (OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS, default 30s), distinct from both timeouts already recorded in omniroute-account-semaphore-timeout.md and omniroute-non-ping-sse-stream-timeout.md. It exists specifically to recover from a stale pooled TCP socket by retrying once on a brand-new connection — but in the incident analyzed here, both the original attempt and the fresh-socket retry timed out, repeatedly, for 46 requests over 48 straight minutes with zero successes. That pattern rules out a stale-socket explanation (a fresh socket bypasses the pool entirely) and points instead at the upstream itself — llama-server, or the R9700 GPU underneath it — being genuinely unresponsive for the whole window. The best primary-source match for that symptom is an open, still-unresolved AMD ROCm bug specific to this exact GPU (ROCm/legacy-rocm-build#6630): an MES-firmware hang during generation on gfx1201/R9700 that leaves the process alive but stuck, sometimes for no logged reason at all. No config change fixes this — raising the 30s timeout only makes each failed attempt take longer to give up, it doesn't un-wedge a hung GPU.

The evidence

Source: omniroute-request-logs-6h-2026-09-15.json, a 339-entry OmniRoute request-log export the user pulled from the dashboard, covering 2026-09-15T12:49:52Z18:44:54Z. Every entry with a non-200 status (66 of them) is a POST /v1/chat/completions against qwen3.8-27b-local (/models/Qwen3.8-27B-UD-Q4_K_XL.gguf), all on the same connectionId (649a2d3e-7527-488e-9b8a-dc4ac2624176) and the same provider (openai-compatible-chat-a7bda643-6687-41f4-b75d-fd2cab746874) — a single upstream connection, not a fan-out artifact.

Sorting every error by timestamp shows one continuous outage, not scattered slow requests:

  • Last successful /v1/chat/completions before the outage: 15:18:12.285Z
  • First failure: 15:25:06.415Z — status 504, "[504]: Direct response did not start within 30000ms — retrying on a fresh socket", duration 60186ms
  • Every single /v1/chat/completions attempt from 15:25:06.415Z through 16:13:29.502Z failed — 46× 504 (all clustered 60021-60324ms, i.e. two back-to-back 30s attempts, both failing) interleaved with 20× 499 ("Request aborted" / "Client disconnected: request_signal_aborted", durations 1.4s-99.97s — these are qwen-code giving up client-side while OmniRoute was still mid-retry)
  • First successful recovery: 16:13:50.388Z, 20873ms — 21 minutes after the last failure attempt cluster, i.e. the very next attempt after the outage window succeeded normally
  • No successful /v1/chat/completions call appears anywhere inside the 15:25:06Z-16:13:29Z window — confirmed by filtering all 154 /v1/chat/completions log entries in that range: every one is 504 or 499.

git log --since=2026-09-14 --until=2026-09-16 shows zero commits in this repo on 2026-09-15 — the outage correlates with no deploy, scripts/update.sh run, or config change on this end.

A second, independent data point (see caveat below): the user also pasted a large raw text table, copied directly from OmniRoute's dashboard UI rather than the JSON export, showing the qwen-classifier connection (qwen3-4b, both the -UD-Q4_K_XL.gguf file this repo's docker-compose.yml currently defaults to, and a -UD-Q8_K_XL.gguf variant that appears nowhere in this repo's checked-in docker-compose.yml/.env.example — either tested by hand against LLAMA_CLASSIFIER_MODEL_FILE outside version control, or evidence of drift worth checking directly on the server) failing repeatedly across two accounts (Haylan, qwen-cli-main), with the same signature: TI: 0|TO: 0 (zero tokens either direction — failed before generating anything) and durations clustering at exactly 60.0-60.4s for the 504s. That's the identical two-attempts-at-30s-each shape as the 27B outage above, strongly suggesting the same underlying mechanism, though — important caveat — this table is not present in the 6-hour JSON export and covers a different, longer time range (timestamps back to 01:21 and 23:54 on unspecified dates), so it cannot be directly time-correlated against the 27B outage above. Treat it as corroborating evidence that this failure mode recurs on both local model connections, not as proof they failed at the same moment.

Root cause: OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS, a mechanism built for a different problem

Confirmed directly against OmniRoute's own source, diegosouzapw/OmniRoute (same repo the two existing timeout docs already cite) — open-sse/utils/directResponseStartTimeout.ts:

const DEFAULT_DIRECT_HEADERS_TIMEOUT_MS = 30_000;
const DIRECT_RESPONSE_START_TIMEOUT_CODE = "DIRECT_RESPONSE_START_TIMEOUT";

export function resolveDirectHeadersTimeoutMs(
  env: Record<string, string | undefined> = process.env
): number {
  const raw = env.OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS;
  if (raw == null || raw.trim() === "") return DEFAULT_DIRECT_HEADERS_TIMEOUT_MS;
  ...
}

function createDirectResponseStartTimeout(timeoutMs: number): Error & { code: string } {
  const err = new Error(
    `Direct response did not start within ${timeoutMs}ms — retrying on a fresh socket`
  ) as Error & { code: string };
  ...
}

"Direct" here means direct (no-proxy) egress — confirmed in open-sse/utils/proxyFetch.ts, which routes any connection with no configured upstream HTTP proxy through this path (as opposed to OmniRoute's separate proxy/relay egress paths). Every local connection in this repo (llama-server, qwen-classifier, both reached over the ai-stack Docker network with no proxy) is "direct" — so this timeout mechanism governs every request to either local model, streaming or non-streaming alike, not just non-streaming JSON responses as the name might suggest.

Why it exists (and why it didn't help here)

Two OmniRoute issues, both with dedicated regression tests in the repo, explain the actual design intent:

  • #4252 (tests/unit/proxyfetch-retry-fresh-socket-4252.test.ts): "Undici dispatcher fails on direct provider requests in 502 bursts" — the default direct dispatcher pools keep-alive sockets; some upstreams silently close idle pooled sockets, so the next request reusing one fails with UND_ERR_SOCKET. Fix: retry once on a fresh, no-keep-alive dispatcher (getRetryDispatcher(), a different instance from getDefaultDispatcher()) so the retry can't grab another already-dead pooled socket.
  • #10214 (tests/unit/proxyfetch-direct-response-start-timeout-10214.test.ts): "Direct (no-proxy) requests stall on a silently-dropped pooled keep-alive socket until the caller's deadline or a service restart" — the harder case: a pooled socket that dies without even an error, just silence. Undici's headersTimeout default (600s) is far too slow to catch this in practice, and the existing #4252 retry never fires because no error is thrown to trigger it. The fix bounds each direct attempt's response-start wait to OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS (30s default) via directFetchWithBoundedResponseStart, and retries once on the same fresh no-keep-alive dispatcher from #4252 when that bound is hit.

Both fixes assume the socket is the problem, not the upstream. directFetchWithBoundedResponseStart's own implementation (open-sse/utils/directResponseStartTimeout.ts) is exactly two attempts: pooled, then fresh. When attempt 2 — a brand-new socket that cannot possibly be a zombie pooled connection — also times out at 30s, the retry logic has nothing left to try and the request fails with DIRECT_RESPONSE_START_TIMEOUT_CODE (surfaced as the 504 seen in the logs). A fresh socket succeeding to connect but the server never sending a response is exactly what "the upstream process is alive but stuck" looks like from OmniRoute's side — it can't distinguish "GPU is wedged mid-generation" from "stale pooled socket," because both present as "nothing came back in 30s." 46 consecutive both-attempts-failed cycles over 48 minutes is far outside what a transient stale-socket burst (the scenario #4252/#10214 were built for) would produce; it's consistent with a sustained upstream outage instead.

Not currently configured in this repo: grep-ing docker-compose.yml and .env.example for OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS finds nothing — this deployment runs on the unmodified 30s default. (This is also, notably, a fourth data point alongside the two already-documented mechanisms and the requestQueue.maxWaitMs/ RATE_LIMIT_EXECUTION_TIMEOUT timer mentioned in passing in omniroute-account-semaphore-timeout.md — OmniRoute has at least four independent timeout knobs guarding different stages of a request's life, three of them 30-second-flavored by default, which is worth keeping in mind the next time an unfamiliar timeout string shows up.)

Why the upstream itself was likely unresponsive: ROCm/legacy-rocm-build#6630

This session has no SSH/shell access to the actual R9700 server — there's no SSH config, and nothing in scripts/ does remote exec, confirmed by inspecting scripts/update.sh and the absence of any ~/.ssh/config entry for the box. So none of the following is confirmed against this specific incident's dmesg/rocm-smi/ docker logs output — it's the closest primary-source match to the symptom, not a diagnosis of this outage.

ROCm/legacy-rocm-build#6630 ("gfx1201 R9700 ROCm 7.14 llama.cpp generation hang, MES queue failure and PSP reset -62; Vulkan passes") is an open, actively investigated issue (created 2026-08-19, most recent update 2026-08-28, no fix landed) that reproduces on the exact same GPU this repo runs on — Radeon AI PRO R9700, gfx1201 — running llama.cpp with -fa on (this repo's llama-server also runs --flash-attn on), with a controlled Vulkan-vs-ROCm A/B: Vulkan completes normally, ROCm hangs during token generation. Direct quotes:

"During the ROCm generation stall: GPU busy reached 100%, memory busy remained 0%, VRAM use was only about 1.1 GB, the container remained alive but made no output progress."

"when MES stops responding, the driver can stay unaware of it indefinitely — the failure only surfaces when something happens to send the next MES message... if a run is left alone after it stops making progress, the kernel prints nothing at all, so a hang can look like a slow workload rather than a fault."

"the failure is probabilistic, not deterministic... a single passing run on this host does not indicate a healthy configuration."

The thread (12 comments as of this research pass, an AMD engineer harkgill-amd participating) has ruled out, one at a time, uni_mes=0, mes_log_enable=1, ROCm 6.4.4, ROCm 7.14, ROCm 10.0.0 stable, the latest TheRock nightly, and GFXOFF-disable — no confirmed fix or workaround exists in the thread as of this research pass. A related comment on the same issue (chrisfranson) reports the identical MES REMOVE_QUEUE/MODE1-reset signature from a completely unrelated workload (headless LibreOffice with OpenCL) on the same gfx1201 silicon, reinforcing that this is a driver/firmware-level fault under general GPU load, not something specific to llama.cpp's request pattern.

This is a different bug from the one already mitigated in this repo. docs/research/rocm-gpu-pin-and-render-group.md already documents and works around ROCm/ROCm#5706 (clock/power pinned at boost whenever two concurrent HIP contexts share the GPU — fixed via GPU_MAX_HW_QUEUES=1, already set on both llama-server and qwen-classifier in docker-compose.yml). #5706's symptom is elevated power draw with the GPU still working; #6630's symptom is generation fully halting with gpu_busy=100%/mem_busy=0% and MES no longer responding at all — a real hang, not a clock-pin inefficiency. GPU_MAX_HW_QUEUES=1 targets #5706's specific trigger (hardware-queue oversubscription across concurrent HIP processes) and has no evidence in #6630's thread of affecting that bug — #6630 reproduces in single-GPU, single-process benchmarks with no second HIP context involved at all, so the already-applied fix should not be assumed to help here.

What would actually resolve this vs. what wouldn't

  • Raising OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS — not recommended as a fix. It would make each failed attempt take longer before giving up (worse latency during a real hang), without addressing why the GPU stopped responding. It's the right lever only if future evidence shows genuinely-slow-but-working responses being mistaken for hangs (the same shape as the already-fixed REQUEST_TIMEOUT_MS issue in omniroute-non-ping-sse-stream-timeout.md) — this incident's 46-for-46 both-attempts-failed pattern over 48 minutes doesn't fit that shape.
  • Concrete next step for whoever has server access when this recurs: check dmesg | grep -i amdgpu and journalctl -k on the R9700 host for MES(...) failed to respond, GPU reset begin, or PSP resume failed lines matching #6630's signature, and docker logs llama-server/docker logs qwen-classifier to see whether the process was alive-but-stuck (consistent with #6630) versus crashed/restarted (which would point elsewhere). Capturing this during a live incident is the only way to move this from "best primary-source match" to "confirmed root cause."
  • Monitor ROCm/legacy-rocm-build#6630 for a fix — it's open and active (AMD engineer engaged as of 2026-08-26); no released ROCm version as of this research pass is confirmed clean.
  • The -UD-Q8_K_XL.gguf classifier variant in the pasted table but absent from version control is worth a direct look on the server (cat .env / docker inspect qwen-classifier for the actual LLAMA_CLASSIFIER_MODEL_FILE in effect) — outside this research pass's reach without server access, flagged here so it isn't lost.

Recovery: how the hang actually clears (or doesn't) — addendum, 2026-09-15

Follow-up question: what actually recovers the socket once this hits, given it's been observed to stay wedged for days at a time? Pulled the full comment thread on ROCm/legacy-rocm-build#6630 directly via the GitHub API (12 comments, angelhalo as primary reporter, harkgill-amd as the responding AMD engineer, plus one corroborating report from chrisfranson on unrelated hardware/workload) — the earlier research pass's source list cited this issue but hadn't read the full thread. Three things fall directly out of it:

There is no reliable in-band recovery. The driver doesn't notice the hang on its own — direct quote: "when MES stops responding, the driver can stay unaware of it indefinitely — the failure only surfaces when something happens to send the next MES message." In a captured live hang, "every driver-managed ring is completely idle while the GPU reports 100% busy," dmesg has zero amdgpu lines, and no task is in D-state — so from the OS's perspective nothing is wrong; only sending the GPU another command (which killing/restarting the stuck process does) triggers the driver to discover the wedge and attempt its own MODE1 reset.

Once triggered, that reset itself is a coin flip across three documented outcomes, not a guaranteed fix:

  1. Clean recoveryGPU reset succeeded, trying to resume, PSP resumes, the card comes back (VRAM is wiped — "VRAM is lost due to GPU reset!" — so the container needs a real restart to reload the model, a plain process respawn isn't enough even when the reset itself works).
  2. Failed resumePSP resume failed, GPU reset end with ret = -62 (the original report's own outcome) — the reset attempt itself fails, leaving the GPU in a worse state than before.
  3. Full kernel soft-lockupchrisfranson's independent report (different workload — headless LibreOffice OpenCL, different card — RX 9070 XT, same gfx1201 silicon) hit outcome 2 or 3 twice out of three times: "the whole system hard-locked (kernel soft lockup pegging a CPU at ~90-100% softirq, requiring a physical power cycle)."

angelhalo deliberately left one hang untouched rather than killing the process, to observe it without contaminating the state with a reset: "I recovered only with a subsequent cold power cycle" — no reset was ever triggered because nothing sent the GPU another message. Their standard test procedure between every single run in this thread is "a cold power cycle (AC removed, ≥30 s)," specifically not a warm/soft reboot — stated reason: "on this card a MODE1 reset takes the host down with it," meaning even the OS's own reboot path can't be trusted to come back cleanly once this GPU is in a bad state. This is the practical answer to "why does it stay stuck for days": nothing about the hang self-clears, docker's restart: unless-stopped policy never fires because the container process is alive and never exits (confirmed: this repo's llama-server and qwen-classifier services have no healthcheck block at all — only omniroute itself does, a plain TCP connect check on its own dashboard port, which says nothing about whether llama-server/qwen-classifier are responding) — so a hang persists until a human notices the symptom (requests failing) and manually intervenes, and "days" is just however long that takes to notice on a homelab box, not a property of the hang itself.

No fix or reliable mitigation exists as of this reading (2026-08-28, the thread's latest comment). harkgill-amd (AMD) could not reproduce locally and asked for a nightly-driver retest; angelhalo retested and it still failed. Every other variable tested still hangs: ROCm 6.4.4 through 10.0.0 stable, TheRock nightlies, amdgpu.uni_mes=0, cwsr_enable=0, mes_log_enable=1, GFXOFF disabled, two different physical R9700 cards, both llama.cpp and vLLM. The thread's own conclusion, as of the last comment: "a probabilistic lost-completion event" with no known trigger to avoid and no known driver/firmware combination that's clean.

Practical takeaway for this repo, given no upstream fix exists:

  • A restart might recover it, might make it worse (failed PSP resume), and might take the whole host down requiring a physical power cycle — there's no way to know in advance which outcome a given hang will produce.
  • Nothing currently watches for this automatically. Docker's restart: unless-stopped is the wrong tool (process doesn't exit) — recovering automatically would need a healthcheck against llama-server's own /health endpoint (llama.cpp's built-in liveness endpoint) paired with something that acts on an unhealthy status, since Docker itself doesn't restart on failed healthchecks without an external watcher (e.g. willfarrell/autoheal or equivalent) — not evaluated here, flagged as a real gap, not a recommendation to implement blind: an automated restart during a hang that's about to fail its PSP resume and lock the host could turn a "requests are failing" incident into "the box needs a physical power cycle" automatically and unattended, which is a real downside worth weighing against faster detection.
  • Given the reset outcome is unpredictable, the safest manual recovery when this is caught live is: restart the affected container, then immediately check dmesg | grep -i amdgpu for PSP resume failed or a soft-lockup signature before assuming it's fixed — if either appears, a full reboot (and per this thread's own testing practice, possibly a genuine AC power cycle rather than a warm reboot) is the next step, not a second restart attempt.

Caveats and open questions

  • JSON export vs. pasted table are two different, non-overlapping captures. The JSON file is a precise 6-hour window with full per-request detail; the pasted table is a longer, dashboard-UI-copied range with less structure and no verifiable overlap with the JSON file's timestamps. A fresh multi-day JSON export (same request-logs endpoint used to produce the file analyzed here) would let a future pass check whether the classifier's failures and the 27B model's outage are literally simultaneous (strong evidence for a shared GPU-level cause) or independent recurrences of the same mechanism on separate schedules.
  • Why the outage self-recovered after ~48 minutes with no observed restart is unexplained. #6630's thread describes hangs resolving via an explicit GPU reset (sometimes failing, requiring reboot) — not a case of a hang clearing on its own after a fixed interval. Nothing in the available data (no server access) confirms whether a restart happened that isn't visible from OmniRoute's logs, or whether this specific hang genuinely self-cleared, which would be a data point against the #6630 hypothesis worth capturing next time.
  • Live reproduction was not attempted. The user suggested testing tool-calls against the classifier via the Windows-side qwen-code CLI (C:\Users\aerli\AppData\Local\qwen-code\bin\qwen.cmd) to try to reproduce a "Direct response did not start" failure live. Skipped for this pass: qwen-code requires an interactive/ already-authenticated session to drive meaningfully, and deliberately trying to reproduce a GPU hang against the shared production classifier risked a genuine 60s+ stall on infrastructure other work depends on, for uncertain diagnostic payoff given the strength of the log-based and source-based evidence already gathered. Worth doing deliberately, with server access on hand to capture rocm-smi/dmesg simultaneously, rather than as a quick check from this pass.

Sources

  • omniroute-request-logs-6h-2026-09-15.json — OmniRoute dashboard request-log export provided by the user (2026-09-15, 339 entries, 12:49:52Z-18:44:54Z)
  • User-pasted OmniRoute dashboard table (qwen-classifier/qwen3-4b failures, separate capture window)
  • diegosouzapw/OmniRouteopen-sse/utils/directResponseStartTimeout.ts, open-sse/utils/proxyFetch.ts, open-sse/utils/proxyDispatcher.ts, tests/unit/proxyfetch-direct-response-start-timeout-10214.test.ts, tests/unit/proxyfetch-retry-fresh-socket-4252.test.ts, open-sse/handlers/chatCore/upstreamTimeouts.ts
  • ROCm/legacy-rocm-build#6630 — open R9700/gfx1201 llama.cpp generation-hang issue, full comment thread (2026-08-19 through 2026-08-28)
  • docs/research/omniroute-account-semaphore-timeout.md — the first already-documented 30s OmniRoute timeout (account semaphore, 429, hardcoded)
  • docs/research/omniroute-non-ping-sse-stream-timeout.md — the second already-documented timeout (first-SSE-event deadline, REQUEST_TIMEOUT_MS-derived)
  • docs/research/rocm-gpu-pin-and-render-group.md — the already-mitigated, different R9700/gfx1201 MES bug (ROCm/ROCm#5706, clock-pin/power, not a hang)
  • Local docker-compose.yml, .env.example (grepped directly, confirming OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS is unset / on the 30s default, and that GPU_MAX_HW_QUEUES=1 is already applied to both GPU services)
  • git log --since=2026-09-14 --until=2026-09-16 (this repo, confirming zero commits during the outage window)

Confidence / uncertainty summary

  • High confidence: the exact mechanism and semantics of OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS / directFetchWithBoundedResponseStart (read directly from OmniRoute's own source and its two regression test files, which spell out the intent in comments referencing the originating issues); the JSON log's chronology, single-connection scope, and the "two ~30s attempts, both failing" shape of every 504 (computed directly from the log file); that this timeout is unconfigured in this repo (direct grep); that no commits landed in this repo during the outage window (direct git log).
  • Medium confidence: that ROCm/legacy-rocm-build#6630 is the actual root cause of this specific outage. The GPU model, driver-family symptom shape (gpu_busy=100%/mem_busy=0%, alive-but-stuck, sometimes logged/sometimes silent), and -fa on usage all match closely, and it's an open/unresolved/actively-discussed issue as of this research pass — but nothing from this incident's own dmesg/rocm-smi output was available to confirm it directly (no SSH access), so this is the best primary-source match to the symptom, not a confirmed diagnosis.
  • Low confidence / open: why the outage recovered on its own after ~48 minutes with no observed restart; whether the classifier's pasted-table failures share the exact same triggering event as the 27B outage analyzed here (same failure signature, but no verified time-overlap between the two data sources); the provenance of the -UD-Q8_K_XL.gguf classifier variant seen in the pasted table but absent from version control.

Live test, 2026-09-15: the classifier is measurably too slow at its own documented worst case — independent of any hang

Before scoping a fix, tested the live qwen-classifier backend directly against realistic worst-case load, per the user's request to gather fresh evidence rather than design blind. Two attempts to reproduce this through qwen-code itself first surfaced an unrelated, separately-useful finding; the direct backend test below is what actually answered the question.

qwen-code's own headless mode never reaches the classifier

Ran qwen --approval-mode auto <prompt> (positional/one-shot, non-interactive) from the Windows-side install (C:\Users\aerli\AppData\Local\qwen-code\bin\qwen.cmd, which has both fastModel and the omniroute-search MCP server already configured), asking it to run a shell dir and use the web-search MCP tool. Both attempts hit a wall before any classifier request was even sent:

  • The MCP tool call was refused outright: Warning: Tool "mcp__omniroute-search__search" requires user approval but cannot execute in non-interactive mode. ... use the -y flag (YOLO mode).
  • The shell tool: the model itself reported run_shell_command as "not registered" in this session and silently substituted a read-only glob call instead — no approval prompt, no classifier call, no system warning printed (unlike the MCP case), across two separate clean runs.

Conclusion: one-shot headless qwen <prompt> invocations don't exercise Auto Mode's classifier at all for approval-requiring tools — they're declined or silently rerouted before the classifier ever gets a request. The classifier only fires in a genuinely interactive session, where it substitutes for the human's live approval decision. This wasn't previously documented anywhere in this repo and is worth keeping in mind: headless qwen-code testing is not a valid way to probe classifier behavior, live or otherwise. (Not investigated further: whether qwen serve/--input-format stream-json headless-agent modes behave differently — plausible, since they're built for exactly this kind of automation, but out of scope for this pass.)

Direct backend test: real classifier latency at realistic token counts

Given headless qwen-code couldn't drive this, sent shaped classifier requests straight to http://proxy-ai.home/v1/chat/completions (model qwen3-4b//models/Qwen3-4B-Instruct-2507-UD-Q4_K_XL.gguf, confirmed present in GET /v1/models) using a temporary scoped API key, mimicking the {shouldBlock} JSON-verdict shape and the MAX_TRANSCRIPT_MESSAGES=40 / MAX_HISTORICAL_ACTION_CHARS=4000 structure this repo's own fast-model-choice.md already read out of qwen-code's classifier-transcript.ts source. Five calls, in order:

# Prompt tokens cached_tokens Wall time Notes
1 1,000 3 2.24s Small prompt, genuinely fresh — healthy baseline.
2 51,234 51,233 64.21s First send of a large synthetic worst-case transcript (~40 highly self-similar 4K-char "historical action" blocks). Near-total cache hit reported, yet still the slowest call — see caveat below.
3 51,234 51,233 0.07s Identical repeat of #2. Same chatcmpl-... id as #2 came back — this is OmniRoute short-circuiting an exact-duplicate request via a proxy-level response cache, not fresh inference. Confirms #2/#3's cached_tokens field is not a reliable proxy for wall-clock latency on its own.
4 1,000 3 0.03s Identical repeat of #1 — same id, same response-cache short-circuit.
5 30,958 919 28.66s Fresh, non-repeated worst-case-shaped transcript (different random content, no internal self-similarity to trigger cache effects). Mostly-uncached (919/30,958) — this is the clean data point.

Call #5 is the one to trust: ~31K genuinely-fresh prompt tokens took 28.7 seconds on the current --n-gpu-layers 28 (of 36) / --cache-type-k/v q4_0 / --parallel 1 configuration, with the backend otherwise idle and healthy (no hang in progress). Extrapolating that rate to the repo's own documented worst case (a real 15,116-token call observed live per fast-model-choice.md, and a theoretical ceiling around 40-50K tokens per classifier-transcript.ts's limits) puts a genuine worst-case classifier call at roughly 30-65 seconds of normal, non-hung processing time — consistent with call #2's 64.21s, even though that call's own cache metadata is too muddied by internal prompt self-similarity to use as a second clean sample.

This directly overlaps both binding timeouts: qwen-code's own client-side classifier stage timeout (stage1Ms/stage2Ms, 60000 each in the Windows-side settings.json observed this session, 30000/60000 in the WSL-side one) and OMNIROUTE_DIRECT_HEADERS_TIMEOUT_MS's 30-second-per-attempt window documented above. A worst-case classifier call landing in the 30-65s range will routinely trip one or both of these timeouts on its own, with the backend never having hung at all — the exact same 499/504/"Request aborted" shape as the GPU-hang hypothesis produces, but from an entirely mundane, deterministic cause: the classifier's current configuration is simply too slow for the request sizes qwen-code's own classifier-transcript design allows.

What this changes

This doesn't rule out ROCm/legacy-rocm-build#6630 — the 27B model's 48-minute total outage (zero successes, not just slow ones) doesn't fit a "just slow" explanation, and remains best matched by a real GPU hang. But it does mean the classifier's own recurring failures (the pasted-table evidence) very plausibly have a second, independent, non-probabilistic cause that a healthcheck/restart-on-hang design wouldn't fix at all — restarting a classifier that's merely slow-but-working at worst-case load just interrupts a call that would have succeeded, and would fire repeatedly under normal peak usage, not just during a rare hang. Any fix that only targets "detect and recover from an unresponsive GPU" leaves this second failure mode untouched. Two independent levers worth weighing before finalizing a scope: raising the classifier's own timeouts to match its real worst-case latency (cheap, immediate, but does nothing for actual hangs), and/or speeding up the classifier itself (full GPU offload if VRAM allows, a faster quant, or capping the transcript size client-side) to bring worst-case latency back under the existing timeouts.

Not investigated in this pass: whether call #2's 64.21s (vs. call #5's extrapolated ~45-48s at a similar token count) reflects genuine non-linear slowdown at the very largest context sizes, real concurrent contention from other production traffic sharing the same --parallel 1 slot during the test, or is just noise from a single sample each — worth a few more clean, uniquely-content, worst-case-sized calls at different times of day before treating either number as precise.