24 KiB
Research: self-hosted alternatives to DashScope for Qwen Code's built-in web_search tool
Question: Qwen Code CLI's built-in web_search tool requires tools.webSearch.model
to resolve to a "DashScope-compatible modelProviders entry." Is there any real,
non-Alibaba-Cloud way to satisfy that requirement with something self-hosted —
or is the already-working OmniRoute MCP + SearXNG path (docs/research/omniroute-qwen-websearch.md)
the end of the road?
Answer, short version: No. The client-side code that decides whether a
baseUrl is "DashScope-compatible" checks the literal hostname against a
hardcoded allowlist of Alibaba-owned domains, before any request is sent — it
is not a protocol-compatibility check that a look-alike server could pass. A
self-hosted server cannot satisfy it, full stop, unless you fork qwen-code and
delete that check. Once you've done that, the actual wire protocol
(OpenAI SDK responses.create(), SSE, specific item types) is buildable
(a few hundred lines), but nothing you can install off the shelf implements it
today. The already-working OmniRoute MCP + SearXNG path costs nothing further
and does not have this problem. Recommendation: don't build this — see
§6.
1. What "DashScope Responses API" is, precisely
Alibaba Cloud Model Studio (Bailian/DashScope) added an OpenAI-compatible Responses API, layered on top of its existing Chat Completions compatible-mode surface:
- Endpoint (per Alibaba's own docs):
POST {baseUrl}/responses, wherebaseUrlis the region's compatible-mode base, e.g.https://dashscope.aliyuncs.com/compatible-mode/v1(China/Beijing) or the-intl/ regional*.maas.aliyuncs.comvariants. Source: https://www.alibabacloud.com/help/en/model-studio/qwen-api-via-openai-responses and https://www.alibabacloud.com/help/en/model-studio/compatibility-with-openai-responses-api - Request shape: standard Responses API (
model,input,stream,store,instructions) plus atoolsarray that can include{"type": "web_search"},{"type": "web_extractor"},{"type": "code_interpreter"}as hosted, server-side tools — the inference backend runs the search itself and streams results back, the same hosted-tool pattern as OpenAI's own Responses APIweb_search_preview, not a client-side function-calling round trip. Source: https://www.alibabacloud.com/help/en/model-studio/qwen-api-via-openai-responses - Response shape: an
outputarray of typed items — aweb_search_callitem carriesaction: {type, query, sources: [{type: "url", url}]}; narration comes back asmessageitems withcontent: [{type, text}]. Source: same page. - Separately, DashScope's plain Chat Completions endpoint (not Responses)
has an older, unrelated
enable_searchboolean (passed viaextra_body) for models like Qwen3.8/Qwen3.6-Plus — the docs explicitly note this older surface does not return citations/sources, which is exactly why qwen-code's built-in tool uses the Responses API instead. Source: https://docs.qwencloud.com/developer-guides/tool-calling/web-search
This is the same hosted-tool pattern as OpenAI's Responses API
web_search_preview (§4 confirms this directly from qwen-code's own client
code — it literally reuses the OpenAI Node SDK's responses.create() call
against a DashScope base URL).
2. What qwen-code's own client code actually sends (ground truth)
Fetched directly from QwenLM/qwen-code's main branch,
packages/core/src/tools/web-search.ts (1087 lines) and
packages/core/src/core/openaiContentGenerator/{constants,provider/dashscope}.ts.
This supersedes anything inferred from the docs pages — it's the literal
implementation.
The request (web-search.ts lines 651–699):
const client = new OpenAI({
apiKey, // from the resolved modelProviders entry's envKey
baseURL: backend.baseUrl, // the modelProviders entry's baseUrl / WEB_SEARCH_BASE_URL
timeout: 60_000,
defaultHeaders: { 'User-Agent': `QwenCode/${version} (...)`, ...customHeaders },
});
const tools = [{ type: 'web_search' }];
if (backend.webExtractor) tools.push({ type: 'web_extractor' });
const requestParams = {
model: backend.modelId,
input: `Perform a web search for the query: ${query}`,
stream: true,
store: false,
instructions: SIDE_REQUEST_INSTRUCTIONS, // a fixed system prompt, see source
tools,
};
const stream = await client.responses.create(requestParams, { signal });
This is the official OpenAI Node SDK, so client.responses.create()
literally POSTs to {baseURL}/responses with that JSON body and reads back
an SSE stream — there is no DashScope-specific SDK involved at all. Anything
speaking real OpenAI Responses-API SSE syntax at that path, with these two
extra tool types, is protocol-compatible.
What the client parses out of the SSE stream (lines 359–509): event types
response.output_item.added, response.output_item.done,
response.output_text.delta, and terminal response.completed /
.failed / .incomplete / .cancelled, each carrying a response object
with output: WsOutputItem[] and usage.x_tools.{web_search,web_extractor}.count.
Output items it understands: web_search_call (action.query/action.queries,
action.sources[].url, status), web_extractor_call (urls, goal,
output, status), and message (content[].text). It also defensively
handles a DashScope-specific quirk: request-level failures arrive as a bare
SSE event:error with {code, message, request_id} and no type/error
wrapper — the OpenAI SDK doesn't recognize this shape, so qwen-code parses
it itself (comment: "probe-verified"). Any replacement backend needs to emit
exactly these item/event shapes, or qwen-code's parser silently ignores
unrecognized item types and ultimately reports
WEB_SEARCH_NO_SEARCH_PERFORMED (it treats zero web_search_call items as
"no search happened," with one retry, before failing outright — see lines
883–906).
The hard gate — this is the actual finding. Before any request is sent,
evaluateWebSearchGate() (lines 169–335) validates the resolved baseUrl
through classifyDashScopeBaseUrl() (lines 122–157):
function classifyDashScopeBaseUrl(baseUrl: string): DashScopeBaseUrlIssue | null {
const url = new URL(baseUrl);
if (url.protocol !== 'https:') return 'insecure';
const hostname = url.hostname.toLowerCase();
const suffixes = [
...DASHSCOPE_REGIONAL_HOSTS, // dashscope.aliyuncs.com, dashscope-intl.aliyuncs.com, dashscope-us.aliyuncs.com
'maas.aliyuncs.com',
'alibaba-inc.com',
'aliyun-inc.com',
];
return suffixes.some(s => hostname === s || hostname.endsWith('.' + s)) ? null : 'unknown-host';
}
DASHSCOPE_REGIONAL_HOSTS is defined in
packages/core/src/core/openaiContentGenerator/provider/dashscope.ts as
exactly ['dashscope.aliyuncs.com', 'dashscope-intl.aliyuncs.com', 'dashscope-us.aliyuncs.com'].
This means "DashScope-compatible" is not a protocol claim you can satisfy
by implementing the right JSON shapes — it is a literal hostname allowlist
checked client-side, before the request is even built. A self-hosted server
at http://search.home, https://proxy-ai.home, or any hostname you control
will be rejected with "WebSearch ... is not a DashScope-compatible
endpoint" regardless of what protocol it speaks, unless its hostname ends in
one of dashscope.aliyuncs.com, dashscope-intl.aliyuncs.com,
dashscope-us.aliyuncs.com, *.maas.aliyuncs.com, *.alibaba-inc.com, or
*.aliyun-inc.com — domains Alibaba owns, that you cannot obtain a valid TLS
certificate for. (There's also a separate, unrelated DASHSCOPE_PROXY_BASE_URL
env var used by the main content generator's provider-detection code
(dashscope.ts lines 244–262) for header/cache-control routing through a
corporate proxy — it is not consulted by classifyDashScopeBaseUrl() at all,
so it does not help here either.)
The only way around this specific check is to fork qwen-code and delete or
relax classifyDashScopeBaseUrl() — it's ~15 lines of open-source
TypeScript, so this is not hard code-wise, but it means running a patched
build of the CLI, not configuring the stock release.
3. Any self-hostable server implementing this surface today? — No
Checked the servers this task named:
- vLLM: has a real
/v1/responsesimplementation (https://docs.vllm.ai/en/stable/api/vllm/entrypoints/openai/responses/), and forgpt-ossmodels specifically supports a built-inbrowsertool with a pluggable, MCP-compliant external tool server in place of the default Exa-backed reference implementation (https://vllm.ai/blog/2025-08-05-gpt-oss; https://github.com/vllm-project/recipes/blob/main/OpenAI/GPT-OSS.md). This is the closest existing building block found — but it's gpt-oss/harmony specific (not Qwen), and its tool/event shapes (browser.search,browser.openharmony-channel messages) are not the same as DashScope'sweb_search_call/web_extractor_callitems qwen-code's parser expects, so it is not drop-in — it would need a translation shim in front, at which point you're building the shim anyway and don't need vLLM in the path. - SGLang: Responses API support is unclear/inconsistent per its own issue tracker (https://github.com/sgl-project/sglang/issues/10038) — no usable built-in web-search tool found.
- LiteLLM: does expose
/v1/responses, but has an open bug rejecting theweb_searchtool type outright — "LiteLLM raises a validation error... onlyweb_search_previewis currently allowed" (https://github.com/BerriAI/litellm/issues/14011). Its actual SearXNG integration is the unrelated standalone/v1/searchREST endpoint already documented indocs/research/litellm-searxng-search.md(§1–3 there) — a sibling API to chat/responses, not a Responses-APItools:[{"type":"web_search"}]handler. It doesn't have a DashScope-mode either (https://docs.litellm.ai/docs/providers/dashscope is a plain client wrapper that calls the real dashscope.aliyuncs.com; nothing in it hosts a DashScope-shaped server). - LocalAI / Ollama: no Responses API or DashScope-compatible mode found in searches for either.
- A generic "OpenAI Responses API" self-hosted shim that could be relabeled:
the closest match found,
teabranch/open-responses-server(185 stars, 161 commits, wraps Ollama/vLLM as a Responses API with MCP support), does not implementweb_searchat all — its own roadmap lists "Web search: crawl4ai" as a future item, not shipped (verified live against the repo, 2026-09-05). No other candidate turned up in repeated GitHub searches for "dashscope emulator/mock/fake server" or "responses api web_search self-hosted."
Conclusion for §3: nothing installable off the shelf implements the
DashScope Responses API's web_search/web_extractor hosted-tool surface.
Building it means writing your own small SSE server (see §5 sizing).
4. Is DashScope's shape "OpenAI Responses API + web_search" reused wholesale?
Yes, confirmed directly from source, not inference: qwen-code's client uses
the official openai npm package's client.responses.create() against a
DashScope baseURL (§2 above) — it is not a DashScope-specific SDK or
protocol. OpenAI's own Responses API supports a hosted web_search_preview
tool with a similar output[].type === "web_search_call" item shape
(OpenAI's public Responses API docs, referenced but not independently
re-fetched here since qwen-code's source is authoritative for what it
actually calls). DashScope's extension is the tool name (web_search
rather than web_search_preview — the exact naming mismatch LiteLLM's own
open bug in §3 stumbles on) plus the additional web_extractor tool and the
x_tools usage-accounting field. No existing "OpenAI Responses API shim"
project was found that already emulates web_search_preview/web_search
server-side against a pluggable backend (see §3) — the two hosted-tool
ecosystems (OpenAI's and DashScope's) both currently require literally
calling out to the vendor's own cloud; nobody has open-sourced a
self-hosted stand-in for either.
5. LiteLLM specifically, re-examined against this exact requirement
docs/research/litellm-searxng-search.md already established SearXNG is a
first-class LiteLLM search_provider behind the standalone /v1/search
REST endpoint (its own §1–2). That endpoint is irrelevant to qwen-code's
tools.webSearch.model gate: qwen-code doesn't call an arbitrary search REST
endpoint, it calls POST {baseUrl}/responses on an OpenAI-SDK client
with tools:[{type:"web_search"}], and gates baseUrl on the Alibaba
hostname allowlist in §2. Even ignoring the hostname gate entirely (i.e.
assuming a patched qwen-code build), LiteLLM's /v1/responses route
currently rejects the web_search tool type per the open bug in §3 — so
today, LiteLLM cannot terminate this request even as an internal component of
a custom build. Nothing here changes the litellm-searxng-search.md
recommendation; it remains correct and unrelated to this question.
6. Effort assessment and recommendation
Option A — patch qwen-code + hand-roll a DashScope-Responses-shaped SSE server in front of SearXNG. What it needs, concretely:
- Fork qwen-code, delete/relax
classifyDashScopeBaseUrl()(§2) — trivial, but means building and distributing a patched CLI, and re-patching on every upstream update that touches this file or its surrounding gate logic. - Write a small HTTP server exposing
POST /responsesthat: accepts the exact request shape in §2, calls SearXNG (http://search.home, already reachable perdocs/research/litellm-searxng-search.md'sextra_hostsfinding) for results, and streams back SSE events in the precise sequence qwen-code's parser expects (response.output_item.added/.donewith aweb_search_callitem carryingaction.sources[].url, optionally amessageitem with narrated text, thenresponse.completed). No narration/LLM step is strictly required — an empty or templatedmessagestill satisfies the parser as long as at least one non-failedweb_search_callitem exists (§2's "no-search" check only counts search-call items, not narration quality). Realistically a few hundred lines (Node/Python + SSE), a day or so of work plus debugging the exact event ordering, error-shape (event:errorquirk), andstore/instructionsfields the client sends but doesn't strictly require echoing back. - Register this server's URL as a
modelProvidersentry — except the patched hostname check from step 1 is required for step 3 to pass at all, so steps 1 and 2 are both mandatory, not alternatives. - Maintain the fork indefinitely against upstream qwen-code releases.
Option B — do nothing further. docs/research/omniroute-qwen-websearch.md
already documents a verified, working, fully self-hosted path: OmniRoute's
own omniroute_web_search MCP tool, backed by this stack's SearXNG instance,
confirmed connected (qwen mcp list → Connected) and exercised end-to-end
(POST /v1/search returned real results). This uses qwen-code's documented,
supported, unpatched MCP-server extension point (mcpServers in
settings.json) — no fork, no upstream-drift risk, no protocol shape to
maintain.
Recommendation: do not build Option A. The built-in web_search tool's
"DashScope-compatible" requirement is, by design in qwen-code's own source, a
hostname allowlist for Alibaba's cloud — it is not a compatibility surface
meant to be reimplemented, and no one else has reimplemented it either (§3).
Satisfying it self-hosted requires forking and permanently maintaining a
patch to code whose only purpose is to stop you from doing that. The MCP
path in omniroute-qwen-websearch.md already delivers the same end-user
capability (web search, backed by this stack's own SearXNG, no external
API) through qwen-code's actual supported extension point, with zero ongoing
fork-maintenance burden. There is no functional gap Option A would close that
Option B doesn't already close today.
Open questions / unknowns
- Whether
DASHSCOPE_REGIONAL_HOSTSor the extra suffixes (maas.aliyuncs.com,alibaba-inc.com,aliyun-inc.com) ever change across qwen-code releases — checked only against the currentmainbranch (fetched 2026-09-05); a future release could tighten or loosen this list. - Whether OpenAI's own
web_search_previewResponses-API tool has a publicly documented exact request/response JSON schema identical enough to DashScope'sweb_search/web_extractorpair that a single shim could serve both — not independently verified against OpenAI's own docs in this pass; qwen-code's source (§2) is authoritative for the DashScope side only. - Whether
teabranch/open-responses-server's planned "Web search: crawl4ai" roadmap item, if shipped, would end up emitting DashScope-shapedweb_search_callitems or OpenAI-shapedweb_search_previewones — could become relevant later but is speculative (unshipped) as of this research.
Sources
- https://qwenlm.github.io/qwen-code-docs/en/developers/tools/web-search/ and https://raw.githubusercontent.com/QwenLM/qwen-code/main/docs/developers/tools/web-search.md — current built-in-tool vs. MCP options, settings keys, migration note.
packages/core/src/tools/web-search.ts,packages/core/src/core/openaiContentGenerator/constants.ts,packages/core/src/core/openaiContentGenerator/provider/dashscope.ts— fetched directly fromQwenLM/qwen-code'smainbranch viaraw.githubusercontent.comon 2026-09-05; ground truth for the request shape, SSE parsing, and the hostname gate (§2).- https://github.com/QwenLM/qwen-code/issues/3841 — prior (closed,
"not planned") community proposal for DashScope
enable_searchpassthrough; shows the feature that eventually shipped took a different path (Responses API, not Chat Completionsenable_search). - https://www.alibabacloud.com/help/en/model-studio/qwen-api-via-openai-responses
and https://www.alibabacloud.com/help/en/model-studio/compatibility-with-openai-responses-api
— Alibaba's own Responses API docs: endpoint,
toolsshape,outputitem shape. - https://docs.qwencloud.com/developer-guides/tool-calling/web-search —
the older Chat-Completions
enable_searchmechanism and its no-citations limitation. - https://docs.vllm.ai/en/stable/api/vllm/entrypoints/openai/responses/, https://vllm.ai/blog/2025-08-05-gpt-oss, https://github.com/vllm-project/recipes/blob/main/OpenAI/GPT-OSS.md — vLLM's Responses API and gpt-oss browser-tool/tool-server support.
- https://github.com/sgl-project/sglang/issues/10038 — SGLang Responses API support unclear.
- https://github.com/BerriAI/litellm/issues/14011 — LiteLLM's
/v1/responsesrejects theweb_searchtool type. - https://docs.litellm.ai/docs/providers/dashscope — LiteLLM's DashScope provider is a plain client wrapper, no Responses API, no web_search.
- https://github.com/teabranch/open-responses-server — closest "self-hosted Responses API shim" found; web_search not implemented (roadmap item only), checked live 2026-09-05.
G:\_DEV\repos\LLM-Server\docs\research\omniroute-qwen-websearch.md— the already-working, verified self-hosted alternative this doc is weighed against.G:\_DEV\repos\LLM-Server\docs\research\litellm-searxng-search.md— LiteLLM's actual (unrelated) SearXNG integration, re-confirmed as orthogonal to this question in §5.
Tried it live (2026-09-05) — confirmed empirically, plus one new fact
The user asked to actually run the experiment rather than stop at the analysis above.
What was done (all local to the WSL install, reverted afterward — nothing in this repo or the live OmniRoute instance was left changed):
- Patched the installed CLI file
~/.local/lib/qwen-code/lib/chunks/web-search-K2FMOGS5.jswith a one-line bypass inclassifyDashScopeBaseUrl():if (baseUrl.includes("proxy-ai.home")) return null; - Added a
tools.webSearchblock to~/.qwen/settings.jsonpointingmodel/baseUrlat a newqwen-experiment-websearchmodelProvidersentry using OmniRoute's existinghttp://proxy-ai.home/v1and the already-workingOMNIROUTE_API_KEY. - Ran
qwenwith a prompt forcing use of the built-inweb_searchtool.
Result — the client-side gate bypass worked, confirming the research's
read of classifyDashScopeBaseUrl() was accurate: qwen accepted the OmniRoute
host as "DashScope-compatible" and attempted the tool call. It stopped at an
interactive approval prompt first (expected — headless auto-approve wasn't
attempted, since that flips on unrestricted auto-execution of every tool call
at process privilege, not just this one).
New fact, not visible from static docs alone: a direct curl -X POST http://proxy-ai.home/v1/responses (with a valid key, matching the request
shape qwen would send) returned {"error":{"message":"No active credentials for provider: codex.","type":"authentication_error","code":"invalid_api_key"}}
— not the generic "unknown route" error a nonexistent path returns (verified
earlier in this same research thread against /v1/search-adjacent bogus
paths). So /v1/responses is a real, implemented OmniRoute route, not
merely undocumented — the earlier inference that it didn't exist was wrong;
it exists but is hardcoded to proxy exclusively through a specific provider
connection OmniRoute's catalog calls codex.
codex identified via PROVIDER_REFERENCE.md: id: codex, alias cx,
name "OpenAI Codex", auth type: OAuth — a real, personal
ChatGPT/OpenAI-account connection, not a free/no-auth scraper provider like
several others already connected in this instance (felo-web,
duckduckgo-web, etc.). Checked docs/reference/ENVIRONMENT.md for any
setting to redirect /v1/responses to a different provider — none
exists; there is no responsesProvider or equivalent override.
Why routing isn't configurable, architecturally: OpenAI's Responses API
web_search is a hosted tool — the search executes inside the model
backend's own infrastructure as part of generating the response, not as a
client-visible round trip. Confirmed directly against llama.cpp's own
tools/server docs (github.com/ggml-org/llama.cpp/tree/master/tools/server):
it implements only /v1/chat/completions with client-side tool-calling
(the model emits a tool_call; the client must execute it), has no
/v1/responses endpoint, no hosted-tool execution, and its built-in
--tools are local-only (read_file, grep_search, exec_shell_command,
etc.) — none make outbound HTTP requests. So even with configurable routing,
pointing /v1/responses at the local Qwen model wouldn't work: the upstream
llama-server has nothing that could serve the hosted-tool half of the
contract. Building that would mean OmniRoute (or a custom shim) intercepting
the model's tool-call mid-generation and splicing in a real search — the
same shim work priced out as not-worth-it earlier in this document, now
confirmed to be the only way, not one option among several.
Conclusion holds, sharpened: the dead end isn't just qwen-code's
client-side hostname check anymore — even a fully self-hosted, hostname-gate-bypassed
setup terminates at OmniRoute's codex-only /v1/responses routing, which
itself terminates at needing a real OpenAI/ChatGPT OAuth account, which is
exactly the kind of external paid dependency this whole line of inquiry was
trying to avoid. omniroute_web_search via MCP (already working, already
free, already self-hosted) remains the only path that actually satisfies the
original goal.
Revert: both the CLI patch and the settings.json changes were reverted
after the test — omniroute-search MCP confirmed still Connected via
qwen mcp list afterward. No lasting changes from this experiment.