Files
LLM-Server/docs/research/dashscope-websearch-selfhost-options.md

413 lines
24 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Research: self-hosted alternatives to DashScope for Qwen Code's built-in `web_search` tool
**Question:** Qwen Code CLI's built-in `web_search` tool requires `tools.webSearch.model`
to resolve to a "DashScope-compatible `modelProviders` entry." Is there any real,
non-Alibaba-Cloud way to satisfy that requirement with something self-hosted —
or is the already-working OmniRoute MCP + SearXNG path (`docs/research/omniroute-qwen-websearch.md`)
the end of the road?
**Answer, short version:** No. The client-side code that decides whether a
`baseUrl` is "DashScope-compatible" checks the **literal hostname** against a
hardcoded allowlist of Alibaba-owned domains, before any request is sent — it
is not a protocol-compatibility check that a look-alike server could pass. A
self-hosted server cannot satisfy it, full stop, unless you fork qwen-code and
delete that check. Once you've done that, the actual wire protocol
(OpenAI SDK `responses.create()`, SSE, specific item types) is buildable
(a few hundred lines), but nothing you can install off the shelf implements it
today. The already-working OmniRoute MCP + SearXNG path costs nothing further
and does not have this problem. **Recommendation: don't build this — see
§6.**
## 1. What "DashScope Responses API" is, precisely
Alibaba Cloud Model Studio (Bailian/DashScope) added an **OpenAI-compatible
Responses API**, layered on top of its existing Chat Completions
compatible-mode surface:
- Endpoint (per Alibaba's own docs): `POST {baseUrl}/responses`, where
`baseUrl` is the region's compatible-mode base, e.g.
`https://dashscope.aliyuncs.com/compatible-mode/v1` (China/Beijing) or the
`-intl` / regional `*.maas.aliyuncs.com` variants.
Source: https://www.alibabacloud.com/help/en/model-studio/qwen-api-via-openai-responses
and https://www.alibabacloud.com/help/en/model-studio/compatibility-with-openai-responses-api
- Request shape: standard Responses API (`model`, `input`, `stream`, `store`,
`instructions`) plus a `tools` array that can include
`{"type": "web_search"}`, `{"type": "web_extractor"}`, `{"type": "code_interpreter"}`
as **hosted, server-side tools** — the inference backend runs the search
itself and streams results back, the same hosted-tool pattern as OpenAI's
own Responses API `web_search_preview`, not a client-side function-calling
round trip.
Source: https://www.alibabacloud.com/help/en/model-studio/qwen-api-via-openai-responses
- Response shape: an `output` array of typed items — a `web_search_call` item
carries `action: {type, query, sources: [{type: "url", url}]}`; narration
comes back as `message` items with `content: [{type, text}]`.
Source: same page.
- Separately, DashScope's plain Chat Completions endpoint (not Responses)
has an older, unrelated `enable_search` boolean (passed via `extra_body`)
for models like Qwen3.8/Qwen3.6-Plus — the docs explicitly note this older
surface **does not return citations/sources**, which is exactly why
qwen-code's built-in tool uses the *Responses* API instead.
Source: https://docs.qwencloud.com/developer-guides/tool-calling/web-search
This is the same hosted-tool pattern as OpenAI's Responses API
`web_search_preview` (§4 confirms this directly from qwen-code's own client
code — it literally reuses the OpenAI Node SDK's `responses.create()` call
against a DashScope base URL).
## 2. What qwen-code's own client code actually sends (ground truth)
Fetched directly from `QwenLM/qwen-code`'s `main` branch,
`packages/core/src/tools/web-search.ts` (1087 lines) and
`packages/core/src/core/openaiContentGenerator/{constants,provider/dashscope}.ts`.
This supersedes anything inferred from the docs pages — it's the literal
implementation.
**The request** (`web-search.ts` lines 651699):
```ts
const client = new OpenAI({
apiKey, // from the resolved modelProviders entry's envKey
baseURL: backend.baseUrl, // the modelProviders entry's baseUrl / WEB_SEARCH_BASE_URL
timeout: 60_000,
defaultHeaders: { 'User-Agent': `QwenCode/${version} (...)`, ...customHeaders },
});
const tools = [{ type: 'web_search' }];
if (backend.webExtractor) tools.push({ type: 'web_extractor' });
const requestParams = {
model: backend.modelId,
input: `Perform a web search for the query: ${query}`,
stream: true,
store: false,
instructions: SIDE_REQUEST_INSTRUCTIONS, // a fixed system prompt, see source
tools,
};
const stream = await client.responses.create(requestParams, { signal });
```
This is the **official OpenAI Node SDK**, so `client.responses.create()`
literally POSTs to `{baseURL}/responses` with that JSON body and reads back
an SSE stream — there is no DashScope-specific SDK involved at all. Anything
speaking real OpenAI Responses-API SSE syntax at that path, with these two
extra tool types, is protocol-compatible.
**What the client parses out of the SSE stream** (lines 359509): event types
`response.output_item.added`, `response.output_item.done`,
`response.output_text.delta`, and terminal `response.completed` /
`.failed` / `.incomplete` / `.cancelled`, each carrying a `response` object
with `output: WsOutputItem[]` and `usage.x_tools.{web_search,web_extractor}.count`.
Output items it understands: `web_search_call` (`action.query`/`action.queries`,
`action.sources[].url`, `status`), `web_extractor_call` (`urls`, `goal`,
`output`, `status`), and `message` (`content[].text`). It also defensively
handles a DashScope-specific quirk: **request-level failures arrive as a bare
SSE `event:error` with `{code, message, request_id}` and no `type`/`error`
wrapper** — the OpenAI SDK doesn't recognize this shape, so qwen-code parses
it itself (comment: "probe-verified"). Any replacement backend needs to emit
exactly these item/event shapes, or qwen-code's parser silently ignores
unrecognized item types and ultimately reports
`WEB_SEARCH_NO_SEARCH_PERFORMED` (it treats zero `web_search_call` items as
"no search happened," with one retry, before failing outright — see lines
883906).
**The hard gate — this is the actual finding.** Before any request is sent,
`evaluateWebSearchGate()` (lines 169335) validates the resolved `baseUrl`
through `classifyDashScopeBaseUrl()` (lines 122157):
```ts
function classifyDashScopeBaseUrl(baseUrl: string): DashScopeBaseUrlIssue | null {
const url = new URL(baseUrl);
if (url.protocol !== 'https:') return 'insecure';
const hostname = url.hostname.toLowerCase();
const suffixes = [
...DASHSCOPE_REGIONAL_HOSTS, // dashscope.aliyuncs.com, dashscope-intl.aliyuncs.com, dashscope-us.aliyuncs.com
'maas.aliyuncs.com',
'alibaba-inc.com',
'aliyun-inc.com',
];
return suffixes.some(s => hostname === s || hostname.endsWith('.' + s)) ? null : 'unknown-host';
}
```
`DASHSCOPE_REGIONAL_HOSTS` is defined in
`packages/core/src/core/openaiContentGenerator/provider/dashscope.ts` as
exactly `['dashscope.aliyuncs.com', 'dashscope-intl.aliyuncs.com', 'dashscope-us.aliyuncs.com']`.
**This means "DashScope-compatible" is not a protocol claim you can satisfy
by implementing the right JSON shapes — it is a literal hostname allowlist
checked client-side, before the request is even built.** A self-hosted server
at `http://search.home`, `https://proxy-ai.home`, or any hostname you control
will be rejected with *"WebSearch ... is not a DashScope-compatible
endpoint"* regardless of what protocol it speaks, unless its hostname ends in
one of `dashscope.aliyuncs.com`, `dashscope-intl.aliyuncs.com`,
`dashscope-us.aliyuncs.com`, `*.maas.aliyuncs.com`, `*.alibaba-inc.com`, or
`*.aliyun-inc.com` — domains Alibaba owns, that you cannot obtain a valid TLS
certificate for. (There's also a separate, unrelated `DASHSCOPE_PROXY_BASE_URL`
env var used by the *main* content generator's provider-detection code
(`dashscope.ts` lines 244262) for header/cache-control routing through a
corporate proxy — it is not consulted by `classifyDashScopeBaseUrl()` at all,
so it does not help here either.)
The only way around this specific check is to **fork qwen-code and delete or
relax `classifyDashScopeBaseUrl()`** — it's ~15 lines of open-source
TypeScript, so this is not hard *code-wise*, but it means running a patched
build of the CLI, not configuring the stock release.
## 3. Any self-hostable server implementing this surface today? — No
Checked the servers this task named:
- **vLLM**: has a real `/v1/responses` implementation
(https://docs.vllm.ai/en/stable/api/vllm/entrypoints/openai/responses/), and
for `gpt-oss` models specifically supports a **built-in `browser` tool**
with a pluggable, MCP-compliant external tool server in place of the
default Exa-backed reference implementation
(https://vllm.ai/blog/2025-08-05-gpt-oss;
https://github.com/vllm-project/recipes/blob/main/OpenAI/GPT-OSS.md). This
is the closest existing building block found — but it's gpt-oss/harmony
specific (not Qwen), and its tool/event shapes (`browser.search`,
`browser.open` harmony-channel messages) are **not** the same as DashScope's
`web_search_call`/`web_extractor_call` items qwen-code's parser expects, so
it is not drop-in — it would need a translation shim in front, at which
point you're building the shim anyway and don't need vLLM in the path.
- **SGLang**: Responses API support is unclear/inconsistent per its own
issue tracker (https://github.com/sgl-project/sglang/issues/10038) — no
usable built-in web-search tool found.
- **LiteLLM**: does expose `/v1/responses`, but has an **open bug**
rejecting the `web_search` tool type outright — "LiteLLM raises a
validation error... only `web_search_preview` is currently allowed"
(https://github.com/BerriAI/litellm/issues/14011). Its actual SearXNG
integration is the unrelated standalone `/v1/search` REST endpoint already
documented in `docs/research/litellm-searxng-search.md` (§13 there) — a
sibling API to chat/responses, not a Responses-API `tools:[{"type":"web_search"}]`
handler. It doesn't have a DashScope-mode either
(https://docs.litellm.ai/docs/providers/dashscope is a plain client wrapper
that calls the real dashscope.aliyuncs.com; nothing in it hosts a
DashScope-shaped server).
- **LocalAI / Ollama**: no Responses API or DashScope-compatible mode found
in searches for either.
- **A generic "OpenAI Responses API" self-hosted shim that could be relabeled**:
the closest match found, `teabranch/open-responses-server` (185 stars, 161
commits, wraps Ollama/vLLM as a Responses API with MCP support), **does not
implement `web_search` at all** — its own roadmap lists "Web search: crawl4ai"
as a *future* item, not shipped (verified live against the repo,
2026-09-05). No other candidate turned up in repeated GitHub searches for
"dashscope emulator/mock/fake server" or "responses api web_search
self-hosted."
**Conclusion for §3: nothing installable off the shelf implements the
DashScope Responses API's `web_search`/`web_extractor` hosted-tool surface.**
Building it means writing your own small SSE server (see §5 sizing).
## 4. Is DashScope's shape "OpenAI Responses API + web_search" reused wholesale?
Yes, confirmed directly from source, not inference: qwen-code's client uses
the **official `openai` npm package**'s `client.responses.create()` against a
DashScope `baseURL` (§2 above) — it is not a DashScope-specific SDK or
protocol. OpenAI's own Responses API supports a hosted `web_search_preview`
tool with a similar `output[].type === "web_search_call"` item shape
(OpenAI's public Responses API docs, referenced but not independently
re-fetched here since qwen-code's source is authoritative for what it
actually calls). DashScope's extension is the tool *name* (`web_search`
rather than `web_search_preview` — the exact naming mismatch LiteLLM's own
open bug in §3 stumbles on) plus the additional `web_extractor` tool and the
`x_tools` usage-accounting field. No existing "OpenAI Responses API shim"
project was found that already emulates `web_search_preview`/`web_search`
server-side against a pluggable backend (see §3) — the two hosted-tool
ecosystems (OpenAI's and DashScope's) both currently require literally
calling out to the vendor's own cloud; nobody has open-sourced a
self-hosted stand-in for either.
## 5. LiteLLM specifically, re-examined against this exact requirement
`docs/research/litellm-searxng-search.md` already established SearXNG is a
first-class LiteLLM `search_provider` behind the **standalone** `/v1/search`
REST endpoint (its own §12). That endpoint is irrelevant to qwen-code's
`tools.webSearch.model` gate: qwen-code doesn't call an arbitrary search REST
endpoint, it calls `POST {baseUrl}/responses` on an **OpenAI-SDK client**
with `tools:[{type:"web_search"}]`, and gates `baseUrl` on the Alibaba
hostname allowlist in §2. Even ignoring the hostname gate entirely (i.e.
assuming a patched qwen-code build), LiteLLM's `/v1/responses` route
currently **rejects** the `web_search` tool type per the open bug in §3 — so
today, LiteLLM cannot terminate this request even as an internal component of
a custom build. Nothing here changes the litellm-searxng-search.md
recommendation; it remains correct and unrelated to this question.
## 6. Effort assessment and recommendation
**Option A — patch qwen-code + hand-roll a DashScope-Responses-shaped SSE
server in front of SearXNG.** What it needs, concretely:
1. Fork qwen-code, delete/relax `classifyDashScopeBaseUrl()` (§2) — trivial,
but means building and distributing a patched CLI, and re-patching on every
upstream update that touches this file or its surrounding gate logic.
2. Write a small HTTP server exposing `POST /responses` that: accepts the
exact request shape in §2, calls SearXNG (`http://search.home`, already
reachable per `docs/research/litellm-searxng-search.md`'s `extra_hosts`
finding) for results, and streams back SSE events in the precise sequence
qwen-code's parser expects (`response.output_item.added` /
`.done` with a `web_search_call` item carrying `action.sources[].url`,
optionally a `message` item with narrated text, then
`response.completed`). No narration/LLM step is strictly required — an
empty or templated `message` still satisfies the parser as long as at
least one non-`failed` `web_search_call` item exists (§2's "no-search"
check only counts search-call items, not narration quality).
Realistically a few hundred lines (Node/Python + SSE), a day or so of
work plus debugging the exact event ordering, error-shape (`event:error`
quirk), and `store`/`instructions` fields the client sends but doesn't
strictly require echoing back.
3. Register this server's URL as a `modelProviders` entry — except the
patched hostname check from step 1 is required for step 3 to pass at all,
so steps 1 and 2 are both mandatory, not alternatives.
4. Maintain the fork indefinitely against upstream qwen-code releases.
**Option B — do nothing further.** `docs/research/omniroute-qwen-websearch.md`
already documents a **verified, working, fully self-hosted** path: OmniRoute's
own `omniroute_web_search` MCP tool, backed by this stack's SearXNG instance,
confirmed connected (`qwen mcp list` → Connected) and exercised end-to-end
(`POST /v1/search` returned real results). This uses qwen-code's *documented,
supported, unpatched* MCP-server extension point (`mcpServers` in
`settings.json`) — no fork, no upstream-drift risk, no protocol shape to
maintain.
**Recommendation: do not build Option A.** The built-in `web_search` tool's
"DashScope-compatible" requirement is, by design in qwen-code's own source, a
hostname allowlist for Alibaba's cloud — it is not a compatibility surface
meant to be reimplemented, and no one else has reimplemented it either (§3).
Satisfying it self-hosted requires forking and permanently maintaining a
patch to code whose only purpose is to *stop* you from doing that. The MCP
path in `omniroute-qwen-websearch.md` already delivers the same end-user
capability (web search, backed by this stack's own SearXNG, no external
API) through qwen-code's actual supported extension point, with zero ongoing
fork-maintenance burden. There is no functional gap Option A would close that
Option B doesn't already close today.
## Open questions / unknowns
- Whether `DASHSCOPE_REGIONAL_HOSTS` or the extra suffixes
(`maas.aliyuncs.com`, `alibaba-inc.com`, `aliyun-inc.com`) ever change
across qwen-code releases — checked only against the current `main` branch
(fetched 2026-09-05); a future release could tighten or loosen this list.
- Whether OpenAI's own `web_search_preview` Responses-API tool has a
publicly documented exact request/response JSON schema identical enough to
DashScope's `web_search`/`web_extractor` pair that a single shim could serve
both — not independently verified against OpenAI's own docs in this pass;
qwen-code's source (§2) is authoritative for the DashScope side only.
- Whether `teabranch/open-responses-server`'s planned "Web search: crawl4ai"
roadmap item, if shipped, would end up emitting DashScope-shaped
`web_search_call` items or OpenAI-shaped `web_search_preview` ones — could
become relevant later but is speculative (unshipped) as of this research.
## Sources
- https://qwenlm.github.io/qwen-code-docs/en/developers/tools/web-search/ and
https://raw.githubusercontent.com/QwenLM/qwen-code/main/docs/developers/tools/web-search.md
— current built-in-tool vs. MCP options, settings keys, migration note.
- `packages/core/src/tools/web-search.ts`,
`packages/core/src/core/openaiContentGenerator/constants.ts`,
`packages/core/src/core/openaiContentGenerator/provider/dashscope.ts`
fetched directly from `QwenLM/qwen-code`'s `main` branch via
`raw.githubusercontent.com` on 2026-09-05; ground truth for the request
shape, SSE parsing, and the hostname gate (§2).
- https://github.com/QwenLM/qwen-code/issues/3841 — prior (closed,
"not planned") community proposal for DashScope `enable_search` passthrough;
shows the feature that eventually shipped took a different path (Responses
API, not Chat Completions `enable_search`).
- https://www.alibabacloud.com/help/en/model-studio/qwen-api-via-openai-responses
and https://www.alibabacloud.com/help/en/model-studio/compatibility-with-openai-responses-api
— Alibaba's own Responses API docs: endpoint, `tools` shape, `output` item
shape.
- https://docs.qwencloud.com/developer-guides/tool-calling/web-search —
the older Chat-Completions `enable_search` mechanism and its
no-citations limitation.
- https://docs.vllm.ai/en/stable/api/vllm/entrypoints/openai/responses/,
https://vllm.ai/blog/2025-08-05-gpt-oss,
https://github.com/vllm-project/recipes/blob/main/OpenAI/GPT-OSS.md — vLLM's
Responses API and gpt-oss browser-tool/tool-server support.
- https://github.com/sgl-project/sglang/issues/10038 — SGLang Responses API
support unclear.
- https://github.com/BerriAI/litellm/issues/14011 — LiteLLM's `/v1/responses`
rejects the `web_search` tool type.
- https://docs.litellm.ai/docs/providers/dashscope — LiteLLM's DashScope
provider is a plain client wrapper, no Responses API, no web_search.
- https://github.com/teabranch/open-responses-server — closest
"self-hosted Responses API shim" found; web_search not implemented
(roadmap item only), checked live 2026-09-05.
- `G:\_DEV\repos\LLM-Server\docs\research\omniroute-qwen-websearch.md`
the already-working, verified self-hosted alternative this doc is weighed
against.
- `G:\_DEV\repos\LLM-Server\docs\research\litellm-searxng-search.md`
LiteLLM's actual (unrelated) SearXNG integration, re-confirmed as
orthogonal to this question in §5.
## Tried it live (2026-09-05) — confirmed empirically, plus one new fact
The user asked to actually run the experiment rather than stop at the analysis above.
**What was done** (all local to the WSL install, reverted afterward — nothing in
this repo or the live OmniRoute instance was left changed):
- Patched the installed CLI file
`~/.local/lib/qwen-code/lib/chunks/web-search-K2FMOGS5.js` with a one-line
bypass in `classifyDashScopeBaseUrl()`: `if (baseUrl.includes("proxy-ai.home")) return null;`
- Added a `tools.webSearch` block to `~/.qwen/settings.json` pointing
`model`/`baseUrl` at a new `qwen-experiment-websearch` `modelProviders` entry
using OmniRoute's existing `http://proxy-ai.home/v1` and the already-working
`OMNIROUTE_API_KEY`.
- Ran `qwen` with a prompt forcing use of the built-in `web_search` tool.
**Result — the client-side gate bypass worked**, confirming the research's
read of `classifyDashScopeBaseUrl()` was accurate: qwen accepted the OmniRoute
host as "DashScope-compatible" and attempted the tool call. It stopped at an
interactive approval prompt first (expected — headless auto-approve wasn't
attempted, since that flips on unrestricted auto-execution of every tool call
at process privilege, not just this one).
**New fact, not visible from static docs alone**: a direct `curl -X POST
http://proxy-ai.home/v1/responses` (with a valid key, matching the request
shape qwen would send) returned `{"error":{"message":"No active credentials
for provider: codex.","type":"authentication_error","code":"invalid_api_key"}}`
**not** the generic "unknown route" error a nonexistent path returns (verified
earlier in this same research thread against `/v1/search`-adjacent bogus
paths). So `/v1/responses` **is a real, implemented OmniRoute route**, not
merely undocumented — the earlier inference that it didn't exist was wrong;
it exists but is hardcoded to proxy exclusively through a specific provider
connection OmniRoute's catalog calls `codex`.
**`codex` identified via `PROVIDER_REFERENCE.md`**: `id: codex`, alias `cx`,
name "OpenAI Codex", **auth type: OAuth** — a real, personal
ChatGPT/OpenAI-account connection, not a free/no-auth scraper provider like
several others already connected in this instance (`felo-web`,
`duckduckgo-web`, etc.). Checked `docs/reference/ENVIRONMENT.md` for any
setting to redirect `/v1/responses` to a different provider — **none
exists**; there is no `responsesProvider` or equivalent override.
**Why routing isn't configurable, architecturally**: OpenAI's Responses API
`web_search` is a *hosted* tool — the search executes inside the model
backend's own infrastructure as part of generating the response, not as a
client-visible round trip. Confirmed directly against llama.cpp's own
`tools/server` docs (`github.com/ggml-org/llama.cpp/tree/master/tools/server`):
it implements only `/v1/chat/completions` with client-side tool-calling
(the model emits a `tool_call`; the *client* must execute it), has no
`/v1/responses` endpoint, no hosted-tool execution, and its built-in
`--tools` are local-only (`read_file`, `grep_search`, `exec_shell_command`,
etc.) — none make outbound HTTP requests. So even with configurable routing,
pointing `/v1/responses` at the local Qwen model wouldn't work: the upstream
llama-server has nothing that could serve the hosted-tool half of the
contract. Building that would mean OmniRoute (or a custom shim) intercepting
the model's tool-call mid-generation and splicing in a real search — the
same shim work priced out as not-worth-it earlier in this document, now
confirmed to be the *only* way, not one option among several.
**Conclusion holds, sharpened**: the dead end isn't just qwen-code's
client-side hostname check anymore — even a fully self-hosted, hostname-gate-bypassed
setup terminates at OmniRoute's `codex`-only `/v1/responses` routing, which
itself terminates at needing a real OpenAI/ChatGPT OAuth account, which is
exactly the kind of external paid dependency this whole line of inquiry was
trying to avoid. `omniroute_web_search` via MCP (already working, already
free, already self-hosted) remains the only path that actually satisfies the
original goal.
**Revert**: both the CLI patch and the `settings.json` changes were reverted
after the test — `omniroute-search` MCP confirmed still `Connected` via
`qwen mcp list` afterward. No lasting changes from this experiment.