diff --git a/docker-compose.yml b/docker-compose.yml index 87bc738..2c6b3d8 100644 --- a/docker-compose.yml +++ b/docker-compose.yml @@ -79,6 +79,14 @@ services: # carried over from llama-server's comment above — same ROCm/ROCm#5706 # clock-pinning bug applies now that two HIP contexts (this + # llama-server) share the card. + # + # --reasoning off is a no-cost safety net, not a confirmed-needed fix: + # ggml-org/llama.cpp#20809 (closed) documents some server builds + # misdetecting Qwen3-Instruct-2507 models as thinking models, routing + # tool-call output into reasoning_content instead of tool_calls — exactly + # the failure mode that ruled out the 27B model for this role in the + # first place. Whether the current image build still has it was never + # independently confirmed (see docs/research/fast-model-choice.md §4/§6). qwen-classifier: image: ghcr.io/ggml-org/llama.cpp:server-rocm container_name: qwen-classifier @@ -105,6 +113,7 @@ services: --batch-size 512 --ubatch-size 128 --flash-attn on + --reasoning off --cache-type-k q4_0 --cache-type-v q4_0 --jinja diff --git a/docs/coding-cli-setup/index.md b/docs/coding-cli-setup/index.md index bf7c368..8d8a44d 100644 --- a/docs/coding-cli-setup/index.md +++ b/docs/coding-cli-setup/index.md @@ -31,4 +31,4 @@ Both serve the same underlying model — `Qwen3.8-27B-UD-Q4_K_XL.gguf`, register | [OpenCode](opencode.md) | OpenAI Chat Completions | `http://:${OMNIROUTE_PORT:-4000}/v1` | `opencode.json` provider block | | [Qwen Code](qwen-code.md) | OpenAI Chat Completions (2 models: chat + `fastModel`) | `http://:${OMNIROUTE_PORT:-4000}/v1` | `~/.qwen/settings.json` `modelProviders.openai` | -Further reading: `docs/research/qwen3.8-27b-tool-calling.md`, `docs/proxy-key-onboarding.md`. +Further reading: `docs/research/qwen3.8-27b-tool-calling.md`, `docs/proxy-key-onboarding.md`, `docs/research/omniroute-account-semaphore-timeout.md` (a connection that can only handle a few concurrent requests — like `llama-server` or `qwen-classifier` — hits a hardcoded 30s reject once more requests queue up than its `maxConcurrent`, unless configured around it). diff --git a/docs/coding-cli-setup/opencode.md b/docs/coding-cli-setup/opencode.md index 1131280..6ab6049 100644 --- a/docs/coding-cli-setup/opencode.md +++ b/docs/coding-cli-setup/opencode.md @@ -25,7 +25,7 @@ curl -fsSL https://opencode.ai/install | bash "models": { "qwen3.8-27b-local": { "name": "Qwen3.8-27B", - "limit": { "context": 65536, "output": 8192 } + "limit": { "context": 131072, "output": 8192 } } } } diff --git a/docs/coding-cli-setup/qwen-code.md b/docs/coding-cli-setup/qwen-code.md index 4040817..69f8639 100644 --- a/docs/coding-cli-setup/qwen-code.md +++ b/docs/coding-cli-setup/qwen-code.md @@ -2,7 +2,18 @@ [← back to overview](index.md) -Qwen Code speaks plain **OpenAI Chat Completions**, and — unlike the other CLIs — needs *two* models: the main chat model, and a `fastModel` for Auto Mode's action classifier (a separate, always-resident, always-fast instance so classification doesn't queue behind chat prefill; see `docker-compose.yml`'s `llama-server-fast` service and `docs/research/fast-model-choice.md`). Both are registered as separate providers in OmniRoute but reachable through the same gateway URL. Config lives in `~/.qwen/settings.json`: +Qwen Code speaks plain **OpenAI Chat Completions**, and — unlike the other CLIs — needs *two* models: the main chat model, and a `fastModel` for Auto Mode's action classifier. Both are registered as separate providers in OmniRoute but reachable through the same gateway URL. + +## Why a second model exists + +Auto Mode's action classifier (`permissions.autoMode`) is qwen-code's per-tool-call safety gate — it decides whether to auto-approve or block a shell command / tool call before it runs. It was originally aliased onto the main 27B model's own OmniRoute connection. That broke two ways in practice (see `docs/research/fast-model-choice.md` for the model research, and the issue-tracker history for the full incident): + +- **Queued behind heavy work.** Every classification call competed for the main model's 2 GPU slots with whatever real generation was already running, so a classifier check could sit blocked for minutes. +- **CPU-only was tried first and was too slow.** Isolating the classifier onto its own CPU-only llama.cpp instance avoided the GPU queue entirely, but real classification calls (which can carry a non-trivial conversation transcript, not just the bare tool call) blew past OmniRoute's request timeout and retry-looped. + +The fix: a dedicated, GPU-resident `qwen-classifier` service (`docker-compose.yml`) running a small model (`Qwen3-4B-Instruct-2507`) on its own **partial** GPU offload — enough layers on the R9700 to be fast, sized to leave real VRAM headroom next to the 27B model rather than trusting a naive weights+KV estimate (see that service's comment block in `docker-compose.yml` for the actual measured numbers and the two wrong turns — batch-size tuning, then flash-attn — before partial offload turned out to be the real lever). + +## `~/.qwen/settings.json` ```json { @@ -16,12 +27,12 @@ Qwen Code speaks plain **OpenAI Chat Completions**, and — unlike the other CLI "generationConfig": { "contextWindowSize": 131072 } }, { - "id": "", - "name": "qwen3.8-27b-classifier", + "id": "", + "name": "qwen3-4b-classifier", "envKey": "OMNIROUTE_API_KEY", "baseUrl": "http://:${OMNIROUTE_PORT:-4000}/v1", "generationConfig": { - "contextWindowSize": 8192, + "contextWindowSize": 65536, "extra_body": { "chat_template_kwargs": { "enable_thinking": false } } } } @@ -32,15 +43,14 @@ Qwen Code speaks plain **OpenAI Chat Completions**, and — unlike the other CLI "name": "", "baseUrl": "http://:${OMNIROUTE_PORT:-4000}/v1" }, - "fastModel": "" + "fastModel": "" } ``` - `envKey` names the environment variable Qwen Code reads the virtual key from — set `OMNIROUTE_API_KEY=` before launching. Both providers can share one virtual key (as above); split it into two if you want separate usage tracking for chat vs. classifier calls. -- **`contextWindowSize` is per-slot, not `LLAMA_CTX_SIZE` itself** — llama.cpp divides `--ctx-size` across `LLAMA_PARALLEL` concurrent slots, and each request only gets one slot's share (same correction applies to OpenCode's `limit.context`). Compute it per model from `.env`: - - Main model: `LLAMA_CTX_SIZE / LLAMA_PARALLEL` = `262144 / 2` = **131072**. - - Fast model: `LLAMA_FAST_CTX_SIZE / LLAMA_FAST_PARALLEL` = `8192 / 1` = **8192**. Undersizing this one specifically breaks Auto Mode ("Classifier stage 1 unavailable") once `hints.allow`/`softDeny`/`hardDeny` entries and recent-action history push a classifier call past it — see the `LLAMA_FAST_CTX_SIZE` comment in `.env.example` before raising it instead of `LLAMA_FAST_PARALLEL`. -- `enable_thinking: false` on the fast model matters: the fast model file (`Qwen3-4B-Instruct-2507`) is already non-thinking, but this also suppresses `` output on any fast-model swap that isn't, keeping classifier responses parseable. +- **`contextWindowSize` for the main model is per-slot, not `LLAMA_CTX_SIZE` itself** — llama.cpp divides `--ctx-size` across `LLAMA_PARALLEL` concurrent slots, and each request only gets one slot's share (same correction applies to OpenCode's `limit.context`). Compute it from `.env`: `LLAMA_CTX_SIZE / LLAMA_PARALLEL` = `262144 / 2` = **131072**. +- **The classifier's `contextWindowSize` (65536) is not per-slot math** — `qwen-classifier` runs `--parallel 1`, so its whole `--ctx-size` belongs to the one slot. 65536 isn't a guess either: qwen-code's own source hard-caps the classifier transcript (`MAX_TRANSCRIPT_MESSAGES=40`, `MAX_HISTORICAL_ACTION_CHARS=4000`/message in `packages/core/src/permissions/classifier-transcript.ts`) — worst case is ~40-50K tokens, so 65536 gives real margin without wasting VRAM the way the original 131072 (copied from the main model's entry, not an actual qwen-code requirement) would have. +- `enable_thinking: false` on the classifier matters for parseability, though `Qwen3-4B-Instruct-2507` is already architecturally non-thinking (see `fast-model-choice.md` §3) — this is belt-and-suspenders for any future fast-model swap that isn't. - Qwen Code also recognizes `advisorModel`, `visionModel`, `compactionModel`, `imageModel` for other model roles — none are wired up in this stack; only `fastModel` is required. ## Web search via OmniRoute @@ -96,9 +106,11 @@ Register it in `~/.qwen/settings.json`: It reuses the same `OMNIROUTE_API_KEY` env var as the model providers above — the virtual key needs search permission in OmniRoute, not just chat-completions. +**Non-interactive mode (`qwen -p ...`) needs this tool explicitly allow-listed.** MCP tools require interactive confirmation by default; `--approval-mode auto` alone doesn't bypass that for a non-interactive run — pass `--allowed-tools mcp__omniroute-search__search` (or `-y` for full YOLO) alongside `-p`, or the search call never reaches the classifier at all and silently no-ops. Confirmed live: without the allow-list, only the tool calls the CLI's non-interactive gate lets through end up as classifier requests. + ## Auto Mode tuning -Auto Mode's action classifier calls the fast model above — its own request can queue behind other stack traffic before the fast llama-server instance is warm, so the default classifier timeout is worth raising. And since this stack is a single trusted local proxy, it's reasonable to pre-approve requests to it rather than confirm every call: +Auto Mode's action classifier calls the fast model above. Even on the dedicated GPU-resident instance, give it real timeout headroom rather than trusting OmniRoute's default — and since this stack is a single trusted local proxy, it's reasonable to pre-approve requests to it rather than confirm every call: ```json { @@ -111,6 +123,8 @@ Auto Mode's action classifier calls the fast model above — its own request can } ``` -`hints.allow` entries are free-text descriptions the classifier matches against, not exact strings — capped at 150 entries/200 chars each (see the `LLAMA_FAST_CTX_SIZE` note above for why that ceiling matters). +`hints.allow` entries are free-text descriptions the classifier matches against, not exact strings — capped at 150 entries/200 chars each. + +Also set a generous per-connection timeout on the classifier's own OmniRoute provider connection (`providerSpecificData.timeoutMs`, dashboard or `PATCH /api/providers/{id}` — not a `.env` value, see `docs/network-access.md` for reaching the dashboard API). 120000ms is comfortable for the current GPU-resident setup (real measured latency: well under a second for a short check, low seconds for the largest realistic transcript) — this isn't the 20-minute figure the main 27B connection needs, since the classifier isn't competing for a contended GPU slot the way the main model can. Everything else in `~/.qwen/settings.json` (`hooks`, `security.auth`'s underlying tooling, editor prefs) is per-machine, not part of pointing at this stack — don't copy it wholesale between machines. diff --git a/docs/research/fast-model-choice.md b/docs/research/fast-model-choice.md index 4f5bac3..ab53b44 100644 --- a/docs/research/fast-model-choice.md +++ b/docs/research/fast-model-choice.md @@ -195,6 +195,39 @@ comfortably affords the higher-precision quant. - [docs/research/qwen3.8-27b-tool-calling.md](qwen3.8-27b-tool-calling.md) (this repo — cross-referenced for the 27B model's own, still-open, tool-calling parser bugs) +## Implementation note (2026-09-09) — what actually shipped, and why it differs + +The model pick (`Qwen3-4B-Instruct-2507`) held up and is what's deployed. Several sizing assumptions in +this doc didn't survive contact with the real deployment, though — worth recording so the next person +tuning this doesn't re-derive the same corrections from scratch: + +- **Service name is `qwen-classifier`, not `llama-server-fast`** — this doc's proposed name never got + used. There's no `LLAMA_FAST_CTX_SIZE`/`LLAMA_FAST_PARALLEL` in `.env.example` either; the real config + lives inline in `docker-compose.yml`'s `qwen-classifier` command. +- **CPU-only was tried first and rejected** — this doc's VRAM budget analysis (§5) assumed GPU + residency from the start, but the actual rollout path tried CPU-only first (to sidestep VRAM + contention entirely) and found it too slow: real classification calls blew past OmniRoute's request + timeout and retry-looped. Moved to GPU after that, which is what §5's math was for all along. +- **Q4_K_XL weights, not Q8_0** — §5's "~2.4GB headroom" case assumed Q8_0 (4.28GB). In practice, fitting + the classifier onto the R9700 *alongside* the 27B model (not in an assumed-empty 7GB budget) left only + ~6.1GB free VRAM total, and even Q4_K_XL (2.37GB) plus full-context KV cache didn't leave enough real + margin at full GPU offload — see the "measured live" numbers in `docker-compose.yml`'s `qwen-classifier` + comment block. Landed on **partial GPU offload (28/36 layers)** instead of full offload, which is not a + case this doc considered at all. +- **65536 context, not 8192** — §5 sized the context "in the low thousands," reasoning from qwen-code's + two-stage classifier description alone. Directly reading qwen-code's actual source + (`packages/core/src/permissions/classifier-transcript.ts`: `MAX_TRANSCRIPT_MESSAGES=40`, + `MAX_HISTORICAL_ACTION_CHARS=4000`/message) puts the real worst case at ~40-50K tokens — confirmed + live, a real classifier call during testing hit 15,116 prompt tokens. 8192 would have been undersized + for real usage; 65536 gives margin without the original setting.json value (131072, copied from the + main model's entry, not a real qwen-code requirement) wasting VRAM for no reason. +- **§4's `--reasoning off` recommendation was initially missed** in the first deployment pass and added + only once this doc was re-read while writing this note. It's now in `docker-compose.yml`'s + `qwen-classifier` command, per this doc's own "add it regardless, no-cost safety net" reasoning — still + unconfirmed whether the current `ghcr.io/ggml-org/llama.cpp:server-rocm` build actually reproduces + #20809 (nothing in testing so far surfaced `reasoning_content` where `tool_calls` was expected, but + that wasn't specifically probed for either). + ## Confidence/uncertainty summary - **High confidence:** Qwen3-4B-Instruct-2507's non-thinking-only status (direct model-card quote); diff --git a/docs/research/omniroute-account-semaphore-timeout.md b/docs/research/omniroute-account-semaphore-timeout.md new file mode 100644 index 0000000..65246f2 --- /dev/null +++ b/docs/research/omniroute-account-semaphore-timeout.md @@ -0,0 +1,84 @@ +# OmniRoute's per-connection semaphore timeout — hardcoded, not a setting + +**Date:** 2026-09-09 + +Any OmniRoute connection whose upstream can only handle a small, fixed number of concurrent requests +(this repo's `llama-server`/`qwen-classifier`, both effectively single-GPU-slot-limited) can hit a hard +30-second reject once more requests are in flight than the connection's `maxConcurrent` allows — even +though the request would have succeeded fine if it had just waited its turn. This surfaced first as the +`pr-agent`/`CodersPlacePI` 429/504 investigation (see the issue tracker), then again while sizing +`qwen-classifier`. Recorded here so it doesn't have to be re-diagnosed from scratch next time. + +## The error + +``` +{"error":{"message":"Semaphore timeout after 30000ms for :","type":"rate_limit_error","code":"rate_limit_exceeded"}} +``` + +## Root cause (confirmed against OmniRoute's own source, [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute)) + +`open-sse/services/accountSemaphore.ts`: + +```ts +const DEFAULT_TIMEOUT_MS = 30_000; +... +function createSemaphoreTimeoutError(semaphoreKey, timeoutMs) { + const error = new Error(`Semaphore timeout after ${timeoutMs}ms for ${semaphoreKey}`); + error.code = "SEMAPHORE_TIMEOUT"; // classified upstream as HTTP 429 rate_limit_exceeded + return error; +} +``` + +Called from `open-sse/handlers/chatCore.ts`: + +```ts +await acquireAccountSemaphore(accountSemaphoreKey, { + maxConcurrency: accountSemaphoreMaxConcurrency, // = the connection's maxConcurrent + signal: streamController.signal, + // no timeoutMs passed → always falls back to the hardcoded 30_000 default +}) +``` + +This is **not** the same thing as OmniRoute's documented quota-share concurrency gate +(`open-sse/services/combo/quotaShareConcurrency.ts`, key prefix `qsconn:`), which is deliberately +fail-open per its own doc comment ("a saturated queue or timeout proceeds without a slot rather than +ever rejecting a dispatchable request") — that one only matters for quota-share combos. The account +semaphore above is a *different*, always-on gate keyed `provider:connectionId`, has no fail-open path, +and its 30-second timeout is a bare `await` with nothing passed to override it — not exposed via +`/api/resilience`, not an env var, not a dashboard toggle, not documented anywhere in +`docs/reference/ENVIRONMENT.md`. It's a hardcoded constant in vendored code. + +Also **not** the same as `requestQueue.maxWaitMs` (visible via `GET /api/resilience`, this deployment +already has it at `86400000`) — that one bounds a Bottleneck-managed *execution* timer that starts only +after dispatch, surfaces as HTTP 504 `RATE_LIMIT_EXECUTION_TIMEOUT`, and is unrelated to the 429 above. + +## What actually fixes it + +The 30s ceiling itself cannot be raised — no config surface reaches it in the current OmniRoute build. +Two real options: + +1. **Bypass the semaphore, let the upstream's own queue absorb concurrency instead.** + `maxConcurrency == null || maxConcurrency <= 0` fully bypasses `accountSemaphore.ts` (see + `isBypassed()`) — no gate, no 30s timer, requests pass straight through to the upstream. This only + works if the upstream itself queues gracefully with no reject-timeout of its own — confirmed true for + llama.cpp's server (`tools/server/server-queue.cpp` has no queue-wait timeout; excess requests just + wait for a free slot). If you do this, also raise the connection's own + `providerSpecificData.timeoutMs` (bounded 1ms–24h, `MAX_PROVIDER_SPECIFIC_TIMEOUT_MS`) generously — + that's the timer that now matters: "did the upstream return response headers in time," which on + llama.cpp means the full queue-wait-then-generate time, since llama.cpp sends **zero bytes, not even + headers**, while a request sits queued (confirmed in `server-context.cpp`: `res->status = 200` is only + set after the first generated token exists). +2. **Reduce how often more than `maxConcurrent` requests actually stack up** — e.g. the + `pr-agent`/Gitea webhook fix (narrowing the subscribed event list so one PR action doesn't fire 3+ + near-simultaneous AI calls). Doesn't remove the ceiling, just makes it less likely to be hit. + +Applied in this repo: `llama-server`'s OmniRoute connection has `maxConcurrent: null` and +`providerSpecificData.timeoutMs: 1200000` (20 min — matches worst-case 2-slots-busy + queued + own +generation time). `qwen-classifier` uses a much shorter `timeoutMs: 120000` since it isn't +GPU-contended the same way — see `docs/coding-cli-setup/qwen-code.md`. + +## Sources + +- [diegosouzapw/OmniRoute](https://github.com/diegosouzapw/OmniRoute) — `open-sse/services/accountSemaphore.ts`, `open-sse/handlers/chatCore.ts`, `open-sse/services/combo/quotaShareConcurrency.ts`, `open-sse/services/rateLimitManager.ts`, `docs/architecture/RESILIENCE_GUIDE.md`, `docs/reference/ENVIRONMENT.md` +- [ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp) — `tools/server/server-queue.cpp`, `tools/server/server-context.cpp` +- `src/shared/validation/providerSpecificData.ts` (OmniRoute) — `MAX_PROVIDER_SPECIFIC_TIMEOUT_MS` bound