413 lines
24 KiB
Markdown
413 lines
24 KiB
Markdown
# Research: self-hosted alternatives to DashScope for Qwen Code's built-in `web_search` tool
|
||
|
||
**Question:** Qwen Code CLI's built-in `web_search` tool requires `tools.webSearch.model`
|
||
to resolve to a "DashScope-compatible `modelProviders` entry." Is there any real,
|
||
non-Alibaba-Cloud way to satisfy that requirement with something self-hosted —
|
||
or is the already-working OmniRoute MCP + SearXNG path (`docs/research/omniroute-qwen-websearch.md`)
|
||
the end of the road?
|
||
|
||
**Answer, short version:** No. The client-side code that decides whether a
|
||
`baseUrl` is "DashScope-compatible" checks the **literal hostname** against a
|
||
hardcoded allowlist of Alibaba-owned domains, before any request is sent — it
|
||
is not a protocol-compatibility check that a look-alike server could pass. A
|
||
self-hosted server cannot satisfy it, full stop, unless you fork qwen-code and
|
||
delete that check. Once you've done that, the actual wire protocol
|
||
(OpenAI SDK `responses.create()`, SSE, specific item types) is buildable
|
||
(a few hundred lines), but nothing you can install off the shelf implements it
|
||
today. The already-working OmniRoute MCP + SearXNG path costs nothing further
|
||
and does not have this problem. **Recommendation: don't build this — see
|
||
§6.**
|
||
|
||
## 1. What "DashScope Responses API" is, precisely
|
||
|
||
Alibaba Cloud Model Studio (Bailian/DashScope) added an **OpenAI-compatible
|
||
Responses API**, layered on top of its existing Chat Completions
|
||
compatible-mode surface:
|
||
|
||
- Endpoint (per Alibaba's own docs): `POST {baseUrl}/responses`, where
|
||
`baseUrl` is the region's compatible-mode base, e.g.
|
||
`https://dashscope.aliyuncs.com/compatible-mode/v1` (China/Beijing) or the
|
||
`-intl` / regional `*.maas.aliyuncs.com` variants.
|
||
Source: https://www.alibabacloud.com/help/en/model-studio/qwen-api-via-openai-responses
|
||
and https://www.alibabacloud.com/help/en/model-studio/compatibility-with-openai-responses-api
|
||
- Request shape: standard Responses API (`model`, `input`, `stream`, `store`,
|
||
`instructions`) plus a `tools` array that can include
|
||
`{"type": "web_search"}`, `{"type": "web_extractor"}`, `{"type": "code_interpreter"}`
|
||
as **hosted, server-side tools** — the inference backend runs the search
|
||
itself and streams results back, the same hosted-tool pattern as OpenAI's
|
||
own Responses API `web_search_preview`, not a client-side function-calling
|
||
round trip.
|
||
Source: https://www.alibabacloud.com/help/en/model-studio/qwen-api-via-openai-responses
|
||
- Response shape: an `output` array of typed items — a `web_search_call` item
|
||
carries `action: {type, query, sources: [{type: "url", url}]}`; narration
|
||
comes back as `message` items with `content: [{type, text}]`.
|
||
Source: same page.
|
||
- Separately, DashScope's plain Chat Completions endpoint (not Responses)
|
||
has an older, unrelated `enable_search` boolean (passed via `extra_body`)
|
||
for models like Qwen3.8/Qwen3.6-Plus — the docs explicitly note this older
|
||
surface **does not return citations/sources**, which is exactly why
|
||
qwen-code's built-in tool uses the *Responses* API instead.
|
||
Source: https://docs.qwencloud.com/developer-guides/tool-calling/web-search
|
||
|
||
This is the same hosted-tool pattern as OpenAI's Responses API
|
||
`web_search_preview` (§4 confirms this directly from qwen-code's own client
|
||
code — it literally reuses the OpenAI Node SDK's `responses.create()` call
|
||
against a DashScope base URL).
|
||
|
||
## 2. What qwen-code's own client code actually sends (ground truth)
|
||
|
||
Fetched directly from `QwenLM/qwen-code`'s `main` branch,
|
||
`packages/core/src/tools/web-search.ts` (1087 lines) and
|
||
`packages/core/src/core/openaiContentGenerator/{constants,provider/dashscope}.ts`.
|
||
This supersedes anything inferred from the docs pages — it's the literal
|
||
implementation.
|
||
|
||
**The request** (`web-search.ts` lines 651–699):
|
||
|
||
```ts
|
||
const client = new OpenAI({
|
||
apiKey, // from the resolved modelProviders entry's envKey
|
||
baseURL: backend.baseUrl, // the modelProviders entry's baseUrl / WEB_SEARCH_BASE_URL
|
||
timeout: 60_000,
|
||
defaultHeaders: { 'User-Agent': `QwenCode/${version} (...)`, ...customHeaders },
|
||
});
|
||
|
||
const tools = [{ type: 'web_search' }];
|
||
if (backend.webExtractor) tools.push({ type: 'web_extractor' });
|
||
|
||
const requestParams = {
|
||
model: backend.modelId,
|
||
input: `Perform a web search for the query: ${query}`,
|
||
stream: true,
|
||
store: false,
|
||
instructions: SIDE_REQUEST_INSTRUCTIONS, // a fixed system prompt, see source
|
||
tools,
|
||
};
|
||
|
||
const stream = await client.responses.create(requestParams, { signal });
|
||
```
|
||
|
||
This is the **official OpenAI Node SDK**, so `client.responses.create()`
|
||
literally POSTs to `{baseURL}/responses` with that JSON body and reads back
|
||
an SSE stream — there is no DashScope-specific SDK involved at all. Anything
|
||
speaking real OpenAI Responses-API SSE syntax at that path, with these two
|
||
extra tool types, is protocol-compatible.
|
||
|
||
**What the client parses out of the SSE stream** (lines 359–509): event types
|
||
`response.output_item.added`, `response.output_item.done`,
|
||
`response.output_text.delta`, and terminal `response.completed` /
|
||
`.failed` / `.incomplete` / `.cancelled`, each carrying a `response` object
|
||
with `output: WsOutputItem[]` and `usage.x_tools.{web_search,web_extractor}.count`.
|
||
Output items it understands: `web_search_call` (`action.query`/`action.queries`,
|
||
`action.sources[].url`, `status`), `web_extractor_call` (`urls`, `goal`,
|
||
`output`, `status`), and `message` (`content[].text`). It also defensively
|
||
handles a DashScope-specific quirk: **request-level failures arrive as a bare
|
||
SSE `event:error` with `{code, message, request_id}` and no `type`/`error`
|
||
wrapper** — the OpenAI SDK doesn't recognize this shape, so qwen-code parses
|
||
it itself (comment: "probe-verified"). Any replacement backend needs to emit
|
||
exactly these item/event shapes, or qwen-code's parser silently ignores
|
||
unrecognized item types and ultimately reports
|
||
`WEB_SEARCH_NO_SEARCH_PERFORMED` (it treats zero `web_search_call` items as
|
||
"no search happened," with one retry, before failing outright — see lines
|
||
883–906).
|
||
|
||
**The hard gate — this is the actual finding.** Before any request is sent,
|
||
`evaluateWebSearchGate()` (lines 169–335) validates the resolved `baseUrl`
|
||
through `classifyDashScopeBaseUrl()` (lines 122–157):
|
||
|
||
```ts
|
||
function classifyDashScopeBaseUrl(baseUrl: string): DashScopeBaseUrlIssue | null {
|
||
const url = new URL(baseUrl);
|
||
if (url.protocol !== 'https:') return 'insecure';
|
||
const hostname = url.hostname.toLowerCase();
|
||
const suffixes = [
|
||
...DASHSCOPE_REGIONAL_HOSTS, // dashscope.aliyuncs.com, dashscope-intl.aliyuncs.com, dashscope-us.aliyuncs.com
|
||
'maas.aliyuncs.com',
|
||
'alibaba-inc.com',
|
||
'aliyun-inc.com',
|
||
];
|
||
return suffixes.some(s => hostname === s || hostname.endsWith('.' + s)) ? null : 'unknown-host';
|
||
}
|
||
```
|
||
|
||
`DASHSCOPE_REGIONAL_HOSTS` is defined in
|
||
`packages/core/src/core/openaiContentGenerator/provider/dashscope.ts` as
|
||
exactly `['dashscope.aliyuncs.com', 'dashscope-intl.aliyuncs.com', 'dashscope-us.aliyuncs.com']`.
|
||
|
||
**This means "DashScope-compatible" is not a protocol claim you can satisfy
|
||
by implementing the right JSON shapes — it is a literal hostname allowlist
|
||
checked client-side, before the request is even built.** A self-hosted server
|
||
at `http://search.home`, `https://proxy-ai.home`, or any hostname you control
|
||
will be rejected with *"WebSearch ... is not a DashScope-compatible
|
||
endpoint"* regardless of what protocol it speaks, unless its hostname ends in
|
||
one of `dashscope.aliyuncs.com`, `dashscope-intl.aliyuncs.com`,
|
||
`dashscope-us.aliyuncs.com`, `*.maas.aliyuncs.com`, `*.alibaba-inc.com`, or
|
||
`*.aliyun-inc.com` — domains Alibaba owns, that you cannot obtain a valid TLS
|
||
certificate for. (There's also a separate, unrelated `DASHSCOPE_PROXY_BASE_URL`
|
||
env var used by the *main* content generator's provider-detection code
|
||
(`dashscope.ts` lines 244–262) for header/cache-control routing through a
|
||
corporate proxy — it is not consulted by `classifyDashScopeBaseUrl()` at all,
|
||
so it does not help here either.)
|
||
|
||
The only way around this specific check is to **fork qwen-code and delete or
|
||
relax `classifyDashScopeBaseUrl()`** — it's ~15 lines of open-source
|
||
TypeScript, so this is not hard *code-wise*, but it means running a patched
|
||
build of the CLI, not configuring the stock release.
|
||
|
||
## 3. Any self-hostable server implementing this surface today? — No
|
||
|
||
Checked the servers this task named:
|
||
|
||
- **vLLM**: has a real `/v1/responses` implementation
|
||
(https://docs.vllm.ai/en/stable/api/vllm/entrypoints/openai/responses/), and
|
||
for `gpt-oss` models specifically supports a **built-in `browser` tool**
|
||
with a pluggable, MCP-compliant external tool server in place of the
|
||
default Exa-backed reference implementation
|
||
(https://vllm.ai/blog/2025-08-05-gpt-oss;
|
||
https://github.com/vllm-project/recipes/blob/main/OpenAI/GPT-OSS.md). This
|
||
is the closest existing building block found — but it's gpt-oss/harmony
|
||
specific (not Qwen), and its tool/event shapes (`browser.search`,
|
||
`browser.open` harmony-channel messages) are **not** the same as DashScope's
|
||
`web_search_call`/`web_extractor_call` items qwen-code's parser expects, so
|
||
it is not drop-in — it would need a translation shim in front, at which
|
||
point you're building the shim anyway and don't need vLLM in the path.
|
||
- **SGLang**: Responses API support is unclear/inconsistent per its own
|
||
issue tracker (https://github.com/sgl-project/sglang/issues/10038) — no
|
||
usable built-in web-search tool found.
|
||
- **LiteLLM**: does expose `/v1/responses`, but has an **open bug**
|
||
rejecting the `web_search` tool type outright — "LiteLLM raises a
|
||
validation error... only `web_search_preview` is currently allowed"
|
||
(https://github.com/BerriAI/litellm/issues/14011). Its actual SearXNG
|
||
integration is the unrelated standalone `/v1/search` REST endpoint already
|
||
documented in `docs/research/litellm-searxng-search.md` (§1–3 there) — a
|
||
sibling API to chat/responses, not a Responses-API `tools:[{"type":"web_search"}]`
|
||
handler. It doesn't have a DashScope-mode either
|
||
(https://docs.litellm.ai/docs/providers/dashscope is a plain client wrapper
|
||
that calls the real dashscope.aliyuncs.com; nothing in it hosts a
|
||
DashScope-shaped server).
|
||
- **LocalAI / Ollama**: no Responses API or DashScope-compatible mode found
|
||
in searches for either.
|
||
- **A generic "OpenAI Responses API" self-hosted shim that could be relabeled**:
|
||
the closest match found, `teabranch/open-responses-server` (185 stars, 161
|
||
commits, wraps Ollama/vLLM as a Responses API with MCP support), **does not
|
||
implement `web_search` at all** — its own roadmap lists "Web search: crawl4ai"
|
||
as a *future* item, not shipped (verified live against the repo,
|
||
2026-09-05). No other candidate turned up in repeated GitHub searches for
|
||
"dashscope emulator/mock/fake server" or "responses api web_search
|
||
self-hosted."
|
||
|
||
**Conclusion for §3: nothing installable off the shelf implements the
|
||
DashScope Responses API's `web_search`/`web_extractor` hosted-tool surface.**
|
||
Building it means writing your own small SSE server (see §5 sizing).
|
||
|
||
## 4. Is DashScope's shape "OpenAI Responses API + web_search" reused wholesale?
|
||
|
||
Yes, confirmed directly from source, not inference: qwen-code's client uses
|
||
the **official `openai` npm package**'s `client.responses.create()` against a
|
||
DashScope `baseURL` (§2 above) — it is not a DashScope-specific SDK or
|
||
protocol. OpenAI's own Responses API supports a hosted `web_search_preview`
|
||
tool with a similar `output[].type === "web_search_call"` item shape
|
||
(OpenAI's public Responses API docs, referenced but not independently
|
||
re-fetched here since qwen-code's source is authoritative for what it
|
||
actually calls). DashScope's extension is the tool *name* (`web_search`
|
||
rather than `web_search_preview` — the exact naming mismatch LiteLLM's own
|
||
open bug in §3 stumbles on) plus the additional `web_extractor` tool and the
|
||
`x_tools` usage-accounting field. No existing "OpenAI Responses API shim"
|
||
project was found that already emulates `web_search_preview`/`web_search`
|
||
server-side against a pluggable backend (see §3) — the two hosted-tool
|
||
ecosystems (OpenAI's and DashScope's) both currently require literally
|
||
calling out to the vendor's own cloud; nobody has open-sourced a
|
||
self-hosted stand-in for either.
|
||
|
||
## 5. LiteLLM specifically, re-examined against this exact requirement
|
||
|
||
`docs/research/litellm-searxng-search.md` already established SearXNG is a
|
||
first-class LiteLLM `search_provider` behind the **standalone** `/v1/search`
|
||
REST endpoint (its own §1–2). That endpoint is irrelevant to qwen-code's
|
||
`tools.webSearch.model` gate: qwen-code doesn't call an arbitrary search REST
|
||
endpoint, it calls `POST {baseUrl}/responses` on an **OpenAI-SDK client**
|
||
with `tools:[{type:"web_search"}]`, and gates `baseUrl` on the Alibaba
|
||
hostname allowlist in §2. Even ignoring the hostname gate entirely (i.e.
|
||
assuming a patched qwen-code build), LiteLLM's `/v1/responses` route
|
||
currently **rejects** the `web_search` tool type per the open bug in §3 — so
|
||
today, LiteLLM cannot terminate this request even as an internal component of
|
||
a custom build. Nothing here changes the litellm-searxng-search.md
|
||
recommendation; it remains correct and unrelated to this question.
|
||
|
||
## 6. Effort assessment and recommendation
|
||
|
||
**Option A — patch qwen-code + hand-roll a DashScope-Responses-shaped SSE
|
||
server in front of SearXNG.** What it needs, concretely:
|
||
1. Fork qwen-code, delete/relax `classifyDashScopeBaseUrl()` (§2) — trivial,
|
||
but means building and distributing a patched CLI, and re-patching on every
|
||
upstream update that touches this file or its surrounding gate logic.
|
||
2. Write a small HTTP server exposing `POST /responses` that: accepts the
|
||
exact request shape in §2, calls SearXNG (`http://search.home`, already
|
||
reachable per `docs/research/litellm-searxng-search.md`'s `extra_hosts`
|
||
finding) for results, and streams back SSE events in the precise sequence
|
||
qwen-code's parser expects (`response.output_item.added` /
|
||
`.done` with a `web_search_call` item carrying `action.sources[].url`,
|
||
optionally a `message` item with narrated text, then
|
||
`response.completed`). No narration/LLM step is strictly required — an
|
||
empty or templated `message` still satisfies the parser as long as at
|
||
least one non-`failed` `web_search_call` item exists (§2's "no-search"
|
||
check only counts search-call items, not narration quality).
|
||
Realistically a few hundred lines (Node/Python + SSE), a day or so of
|
||
work plus debugging the exact event ordering, error-shape (`event:error`
|
||
quirk), and `store`/`instructions` fields the client sends but doesn't
|
||
strictly require echoing back.
|
||
3. Register this server's URL as a `modelProviders` entry — except the
|
||
patched hostname check from step 1 is required for step 3 to pass at all,
|
||
so steps 1 and 2 are both mandatory, not alternatives.
|
||
4. Maintain the fork indefinitely against upstream qwen-code releases.
|
||
|
||
**Option B — do nothing further.** `docs/research/omniroute-qwen-websearch.md`
|
||
already documents a **verified, working, fully self-hosted** path: OmniRoute's
|
||
own `omniroute_web_search` MCP tool, backed by this stack's SearXNG instance,
|
||
confirmed connected (`qwen mcp list` → Connected) and exercised end-to-end
|
||
(`POST /v1/search` returned real results). This uses qwen-code's *documented,
|
||
supported, unpatched* MCP-server extension point (`mcpServers` in
|
||
`settings.json`) — no fork, no upstream-drift risk, no protocol shape to
|
||
maintain.
|
||
|
||
**Recommendation: do not build Option A.** The built-in `web_search` tool's
|
||
"DashScope-compatible" requirement is, by design in qwen-code's own source, a
|
||
hostname allowlist for Alibaba's cloud — it is not a compatibility surface
|
||
meant to be reimplemented, and no one else has reimplemented it either (§3).
|
||
Satisfying it self-hosted requires forking and permanently maintaining a
|
||
patch to code whose only purpose is to *stop* you from doing that. The MCP
|
||
path in `omniroute-qwen-websearch.md` already delivers the same end-user
|
||
capability (web search, backed by this stack's own SearXNG, no external
|
||
API) through qwen-code's actual supported extension point, with zero ongoing
|
||
fork-maintenance burden. There is no functional gap Option A would close that
|
||
Option B doesn't already close today.
|
||
|
||
## Open questions / unknowns
|
||
|
||
- Whether `DASHSCOPE_REGIONAL_HOSTS` or the extra suffixes
|
||
(`maas.aliyuncs.com`, `alibaba-inc.com`, `aliyun-inc.com`) ever change
|
||
across qwen-code releases — checked only against the current `main` branch
|
||
(fetched 2026-09-05); a future release could tighten or loosen this list.
|
||
- Whether OpenAI's own `web_search_preview` Responses-API tool has a
|
||
publicly documented exact request/response JSON schema identical enough to
|
||
DashScope's `web_search`/`web_extractor` pair that a single shim could serve
|
||
both — not independently verified against OpenAI's own docs in this pass;
|
||
qwen-code's source (§2) is authoritative for the DashScope side only.
|
||
- Whether `teabranch/open-responses-server`'s planned "Web search: crawl4ai"
|
||
roadmap item, if shipped, would end up emitting DashScope-shaped
|
||
`web_search_call` items or OpenAI-shaped `web_search_preview` ones — could
|
||
become relevant later but is speculative (unshipped) as of this research.
|
||
|
||
## Sources
|
||
|
||
- https://qwenlm.github.io/qwen-code-docs/en/developers/tools/web-search/ and
|
||
https://raw.githubusercontent.com/QwenLM/qwen-code/main/docs/developers/tools/web-search.md
|
||
— current built-in-tool vs. MCP options, settings keys, migration note.
|
||
- `packages/core/src/tools/web-search.ts`,
|
||
`packages/core/src/core/openaiContentGenerator/constants.ts`,
|
||
`packages/core/src/core/openaiContentGenerator/provider/dashscope.ts` —
|
||
fetched directly from `QwenLM/qwen-code`'s `main` branch via
|
||
`raw.githubusercontent.com` on 2026-09-05; ground truth for the request
|
||
shape, SSE parsing, and the hostname gate (§2).
|
||
- https://github.com/QwenLM/qwen-code/issues/3841 — prior (closed,
|
||
"not planned") community proposal for DashScope `enable_search` passthrough;
|
||
shows the feature that eventually shipped took a different path (Responses
|
||
API, not Chat Completions `enable_search`).
|
||
- https://www.alibabacloud.com/help/en/model-studio/qwen-api-via-openai-responses
|
||
and https://www.alibabacloud.com/help/en/model-studio/compatibility-with-openai-responses-api
|
||
— Alibaba's own Responses API docs: endpoint, `tools` shape, `output` item
|
||
shape.
|
||
- https://docs.qwencloud.com/developer-guides/tool-calling/web-search —
|
||
the older Chat-Completions `enable_search` mechanism and its
|
||
no-citations limitation.
|
||
- https://docs.vllm.ai/en/stable/api/vllm/entrypoints/openai/responses/,
|
||
https://vllm.ai/blog/2025-08-05-gpt-oss,
|
||
https://github.com/vllm-project/recipes/blob/main/OpenAI/GPT-OSS.md — vLLM's
|
||
Responses API and gpt-oss browser-tool/tool-server support.
|
||
- https://github.com/sgl-project/sglang/issues/10038 — SGLang Responses API
|
||
support unclear.
|
||
- https://github.com/BerriAI/litellm/issues/14011 — LiteLLM's `/v1/responses`
|
||
rejects the `web_search` tool type.
|
||
- https://docs.litellm.ai/docs/providers/dashscope — LiteLLM's DashScope
|
||
provider is a plain client wrapper, no Responses API, no web_search.
|
||
- https://github.com/teabranch/open-responses-server — closest
|
||
"self-hosted Responses API shim" found; web_search not implemented
|
||
(roadmap item only), checked live 2026-09-05.
|
||
- `G:\_DEV\repos\LLM-Server\docs\research\omniroute-qwen-websearch.md` —
|
||
the already-working, verified self-hosted alternative this doc is weighed
|
||
against.
|
||
- `G:\_DEV\repos\LLM-Server\docs\research\litellm-searxng-search.md` —
|
||
LiteLLM's actual (unrelated) SearXNG integration, re-confirmed as
|
||
orthogonal to this question in §5.
|
||
|
||
## Tried it live (2026-09-05) — confirmed empirically, plus one new fact
|
||
|
||
The user asked to actually run the experiment rather than stop at the analysis above.
|
||
|
||
**What was done** (all local to the WSL install, reverted afterward — nothing in
|
||
this repo or the live OmniRoute instance was left changed):
|
||
- Patched the installed CLI file
|
||
`~/.local/lib/qwen-code/lib/chunks/web-search-K2FMOGS5.js` with a one-line
|
||
bypass in `classifyDashScopeBaseUrl()`: `if (baseUrl.includes("proxy-ai.home")) return null;`
|
||
- Added a `tools.webSearch` block to `~/.qwen/settings.json` pointing
|
||
`model`/`baseUrl` at a new `qwen-experiment-websearch` `modelProviders` entry
|
||
using OmniRoute's existing `http://proxy-ai.home/v1` and the already-working
|
||
`OMNIROUTE_API_KEY`.
|
||
- Ran `qwen` with a prompt forcing use of the built-in `web_search` tool.
|
||
|
||
**Result — the client-side gate bypass worked**, confirming the research's
|
||
read of `classifyDashScopeBaseUrl()` was accurate: qwen accepted the OmniRoute
|
||
host as "DashScope-compatible" and attempted the tool call. It stopped at an
|
||
interactive approval prompt first (expected — headless auto-approve wasn't
|
||
attempted, since that flips on unrestricted auto-execution of every tool call
|
||
at process privilege, not just this one).
|
||
|
||
**New fact, not visible from static docs alone**: a direct `curl -X POST
|
||
http://proxy-ai.home/v1/responses` (with a valid key, matching the request
|
||
shape qwen would send) returned `{"error":{"message":"No active credentials
|
||
for provider: codex.","type":"authentication_error","code":"invalid_api_key"}}`
|
||
— **not** the generic "unknown route" error a nonexistent path returns (verified
|
||
earlier in this same research thread against `/v1/search`-adjacent bogus
|
||
paths). So `/v1/responses` **is a real, implemented OmniRoute route**, not
|
||
merely undocumented — the earlier inference that it didn't exist was wrong;
|
||
it exists but is hardcoded to proxy exclusively through a specific provider
|
||
connection OmniRoute's catalog calls `codex`.
|
||
|
||
**`codex` identified via `PROVIDER_REFERENCE.md`**: `id: codex`, alias `cx`,
|
||
name "OpenAI Codex", **auth type: OAuth** — a real, personal
|
||
ChatGPT/OpenAI-account connection, not a free/no-auth scraper provider like
|
||
several others already connected in this instance (`felo-web`,
|
||
`duckduckgo-web`, etc.). Checked `docs/reference/ENVIRONMENT.md` for any
|
||
setting to redirect `/v1/responses` to a different provider — **none
|
||
exists**; there is no `responsesProvider` or equivalent override.
|
||
|
||
**Why routing isn't configurable, architecturally**: OpenAI's Responses API
|
||
`web_search` is a *hosted* tool — the search executes inside the model
|
||
backend's own infrastructure as part of generating the response, not as a
|
||
client-visible round trip. Confirmed directly against llama.cpp's own
|
||
`tools/server` docs (`github.com/ggml-org/llama.cpp/tree/master/tools/server`):
|
||
it implements only `/v1/chat/completions` with client-side tool-calling
|
||
(the model emits a `tool_call`; the *client* must execute it), has no
|
||
`/v1/responses` endpoint, no hosted-tool execution, and its built-in
|
||
`--tools` are local-only (`read_file`, `grep_search`, `exec_shell_command`,
|
||
etc.) — none make outbound HTTP requests. So even with configurable routing,
|
||
pointing `/v1/responses` at the local Qwen model wouldn't work: the upstream
|
||
llama-server has nothing that could serve the hosted-tool half of the
|
||
contract. Building that would mean OmniRoute (or a custom shim) intercepting
|
||
the model's tool-call mid-generation and splicing in a real search — the
|
||
same shim work priced out as not-worth-it earlier in this document, now
|
||
confirmed to be the *only* way, not one option among several.
|
||
|
||
**Conclusion holds, sharpened**: the dead end isn't just qwen-code's
|
||
client-side hostname check anymore — even a fully self-hosted, hostname-gate-bypassed
|
||
setup terminates at OmniRoute's `codex`-only `/v1/responses` routing, which
|
||
itself terminates at needing a real OpenAI/ChatGPT OAuth account, which is
|
||
exactly the kind of external paid dependency this whole line of inquiry was
|
||
trying to avoid. `omniroute_web_search` via MCP (already working, already
|
||
free, already self-hosted) remains the only path that actually satisfies the
|
||
original goal.
|
||
|
||
**Revert**: both the CLI patch and the `settings.json` changes were reverted
|
||
after the test — `omniroute-search` MCP confirmed still `Connected` via
|
||
`qwen mcp list` afterward. No lasting changes from this experiment.
|