- qwen_delegate MCP tool: src/qwen-delegate.ts (testable core, subprocess spawn/timeout/parse) + src/qwen-delegate-server.ts (thin MCP stdio wiring) - test/qwen-delegate.test.ts (node:test, mocked spawn) - docs/agents/* + CLAUDE.md from /setup-matt-pocock-skills (Gitea issue tracker via tea CLI, default triage labels, single-context domain docs) - research/qwen-mcp-delegation.md, corrected after confirming qwen-code runs natively on Windows (no WSL) against a local OpenAI-compatible proxy Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CRnb5Gqdu7gVTQrwAqFdfJ
17 KiB
Delegating "light work" from Claude Code to a local qwen (WSL) harness via MCP
Researched 2026-09-06. Primary sources cited inline; no WSL/qwen commands were executed (none available in this environment) — see Open Questions at the end for what to verify locally.
Update (built, same day): the WSL assumption below was wrong.
qwen-codeis installed natively on Windows (C:\Users\aerli\AppData\Local\qwen-code\bin\qwen.cmd, on PATH asqwen), backed by an OpenAI-compatible endpoint athttp://proxy-ai.home/v1viaOMNIROUTE_API_KEY(already configured in~/.qwen/settings.json). No WSL hop needed — Candidate 4's shell-out plan applies directly on Windows. Built as a Node/TS MCP server in this repo (src/qwen-delegate-server.ts+src/qwen-delegate.ts), registered globally viaclaude mcp add --scope user, with tests intest/qwen-delegate.test.ts. Live latency is real — a trivial one-word prompt took 3-6 minutes end to end — so the tool uses a generous (10 min) timeout rather than trying to enforce "light" in code.
Summary / Recommendation
- Candidate 1 (
llm-wrapper-mcp-server): viable only if you route through OpenRouter's hosted API. It talks HTTP to an OpenAI-style/chat/completions-shaped endpoint (defaulthttps://openrouter.ai/api/v1, overridable viaLLM_API_BASE_URL). It does not shell out to a CLI, and qwen (theqwenCLI/harness you run in WSL) has no HTTP endpoint of its own — so this project doesn't connect to your actual setup unless you first stand up an OpenAI-compatible server in front of qwen (e.g. run the underlying qwen3.8 weights via Ollama/vLLM instead of theqwenCLI). Verdict: not directly applicable to awsl qwen -p "..."CLI harness; only viable if you bypass theqwenCLI and serve the model via Ollama/vLLM. - Candidate 2 (
harness/mcp-server): confirmed dead end — this is Harness.io's CI/CD platform MCP server, unrelated to LLM delegation or "qwen harness." Pure name collision. - Candidate 3 (qwen-code CLI): this is almost certainly what the user means by "qwen harness." It's Apache-2.0, originally forked from Gemini CLI, and does have a genuine headless/scriptable mode (
qwen -p "<prompt>") plus MCP client support and asettings.jsonmodelProvidersmechanism that can point at a local OpenAI-compatible endpoint (Ollama/vLLM) if that's how the qwen3.8 model is actually being served. It is not itself an MCP server for arbitrary delegation (though it has an experimentalqwen serveHTTP+SSE daemon mode for its own ACP sessions, not a general MCP tool endpoint). - Candidate 4 (generic/custom MCP proxy): this is the realistic path. There is no first-party or well-maintained "shell out to a CLI and return stdout as an MCP tool" server in the official
modelcontextprotocol/serversrepo. The pragmatic, low-effort solution is a ~50-line custom MCP server (Python SDK or TypeScript SDK) with one tool (e.g.qwen_delegate) that runswsl.exe qwen -p "<prompt>"(orwsl.exe bash -lc "qwen -p '...'"), captures stdout, and returns it as the tool result. - Candidate 5 (Claude Code routing): Claude Code has no built-in model-routing/delegation config — there's no "route X to tool Y" setting. Routing to an MCP tool for "light work" is purely a prompted convention: you write a rule into
CLAUDE.md(e.g. "for simple lookups/light edits, call theqwen_delegateMCP tool instead of doing it yourself") and Claude Code's own judgment (as an LLM reading its system/project instructions) decides when to invoke the tool. MCP server registration itself (.mcp.json/claude mcp add) is well-documented and directly supports a stdio server whosecommandiswsl.exe.
Recommended concrete plan: build the minimal custom MCP server described in Candidate 4, register it as a project-scoped stdio server in .mcp.json with command: wsl.exe, and add an explicit routing instruction to CLAUDE.md. Do not adopt llm-wrapper-mcp-server unless you decide to serve qwen3.8 via Ollama/vLLM behind an OpenAI-compatible endpoint instead of the qwen CLI.
Candidate 1: matdev83/llm-wrapper-mcp-server
Source: GitHub README, PyPI (PyPI page failed to render for automated fetch/search beyond confirming it exists under this name and license — see Open Questions).
- Backends supported: Primarily OpenRouter.ai. The README's own description: "Allow any MCP-capable LLM agent to communicate with or delegate tasks to any other LLM available through the OpenRouter.ai API." Base URL defaults to
https://openrouter.ai/api/v1and is overridable viaLLM_API_BASE_URLenv var or--llm-api-base-urlCLI flag — so in principle it can point at any OpenAI-compatible HTTP endpoint, but it is an HTTP client, not a CLI-shell-out wrapper. - CLI shell-out: No evidence of any capability to invoke an external CLI process (like
qwen). It only ever makes HTTP requests. - MCP tool surface: One tool,
llm_call(prompt: str, model: str | None). Standard MCP methods (initialize,tools/list,resources/list) via stdio/JSON-RPC. - Configuration: env vars
OPENROUTER_API_KEY(required),LLM_API_BASE_URL(optional); CLI flags--model(defaultperplexity/llama-3.1-sonar-small-128k-online),--llm-api-base-url,--log-level;.envfile support viapython-dotenv. - License: MIT.
- Maintenance: ~42 commits on main, 0 stars/forks/watchers visible on the repo page — very low adoption, no evidence of active community use.
- Install/run:
pip install llm-wrapper-mcp-server, thenpython -m llm_wrapper_mcp_server [OPTIONS]. - Dependencies: pydantic, requests, tiktoken, llm-accounting.
Verdict: not applicable as-is. It solves "delegate to another LLM over HTTP," but qwen as a CLI/WSL harness has no HTTP API to call. It becomes viable only if you additionally stand up an OpenAI-compatible server (Ollama or vLLM) serving the qwen3.8 weights directly, bypassing the qwen CLI/harness entirely — which changes the architecture the user described (they specifically want to drive the qwen CLI harness, which per Candidate 3 has its own search/RAG module wired in that a bare model server wouldn't have).
Candidate 2: harness/mcp-server
Source: GitHub README.
Confirmed dead end / naming collision. This is Harness.io's own MCP server for the Harness CI/CD/DevOps platform: "An MCP (Model Context Protocol) server that gives AI agents full access to the Harness.io platform through 11 consolidated tools and 243 resource types" — pipelines, GitOps, feature flags, cloud cost management, security testing. It has nothing to do with LLM delegation, "qwen harness," or routing prompts to a secondary model. Note explicitly for the user: the word "harness" here refers to the company Harness Inc., not to a "coding harness" wrapping an LLM.
Candidate 3: The "qwen harness" / qwen-code CLI
Source: QwenLM/qwen-code GitHub, Qwen Code Docs — Configuration/Settings, Qwen Code Docs — Model Providers, related issue QwenLM/qwen-code#3384 (OpenAI-compatible local LLM support).
- This is almost certainly the tool the user means by "qwen harness"/"qwen code": a coding-agent CLI, originally forked from Google's Gemini CLI ("This project was originally based on Google Gemini CLI v0.8.2"; the project stopped syncing with upstream after Qwen Code v0.1).
- License: Apache 2.0.
- MCP role: Qwen Code is an MCP client (it can call out to MCP servers itself — README lists "MCP, Plan Mode, LSP Integration" as capabilities). It is not documented as exposing a general-purpose MCP server interface for other agents to call into. There is a separate experimental
qwen servemode described as a "shared agent session over HTTP+SSE (ACP) — multiple clients, one agent," which is its own Agent Communication Protocol daemon, not a standard MCP tool endpoint — do not conflate the two. - Non-interactive/scriptable mode: Yes —
qwen -p "<prompt>"is documented for "Scripts, CI/CD, batch processing — no UI." This is the flag to wrap in a shell-out MCP tool. - Backend flexibility: README states support for "OpenAI, Anthropic, Gemini, and Qwen APIs. Any third-party provider or local model (Ollama / vLLM). Switch at runtime." The Model Providers doc confirms
modelProvidersentries insettings.jsoncan point athttp://localhost:11434/v1(Ollama) orhttp://localhost:8000/v1(vLLM), and acontextWindowSizeoverride exists ingenerationConfigfor providers whose effective limit differs from qwen-code's built-in defaults — relevant since the user's qwen3.8 instance has a ~127k window that may need an explicit override rather than relying on qwen-code's name-based default table.
Verdict: this is the real target to wrap, via its -p non-interactive flag, invoked through wsl.exe qwen -p "..." from Windows. It is not itself pluggable as an MCP server for Claude Code to call directly — you still need a small MCP shim.
Candidate 4: Generic MCP LLM-proxy alternatives / custom server
Sources: modelcontextprotocol/servers, MCP Python SDK, MCP TypeScript SDK.
- The official
modelcontextprotocol/serversrepo has no first-party "generic OpenAI-compatible LLM proxy" or "shell out to CLI" server; it only ships reference servers (filesystem, git, memory, fetch, and the "Everything" test/reference server demonstrating prompts/resources/tools generically, not LLM delegation specifically). - No well-maintained community equivalent surfaced in this repo either — the space is thin, matching the low-star
llm-wrapper-mcp-serverfinding above. - Given qwen has no native OpenAI-compatible HTTP API when run as the
qwenCLI harness (only when the underlying model is separately served via Ollama/vLLM), writing a minimal custom MCP server that shells outwsl qwen -p "<prompt>"and returns stdout is the more realistic path than adapting either of the two candidates the user found. This is a small, single-tool stdio MCP server (using the official Python or TypeScript SDK) — roughly:- one tool, e.g.
qwen_delegate(prompt: string) -> string - implementation: spawn
wsl.exe(orwsl.exe -e bash -lc "qwen -p '...'"if quoting/shell needs matter) as a subprocess, capture stdout/stderr, return stdout (trimmed) as the tool result, with a timeout and basic error surfacing (non-zero exit code, stderr content) since qwen's research/SearXNG module may be slow or occasionally fail. - this keeps qwen's own internet/search capability intact (unlike swapping to a bare Ollama/vLLM completions endpoint, which would lose that SearXNG-backed research tool entirely).
- one tool, e.g.
Candidate 5: Instructing Claude Code to route to the MCP tool
Sources: Claude Code MCP docs (redirected from docs.claude.com/en/docs/claude-code/mcp).
- No built-in routing/delegation config exists. Claude Code does not have a mechanism like "send requests matching pattern X to MCP tool Y automatically." Any "route light work to qwen" behavior is a prompted convention: you instruct Claude Code (via
CLAUDE.md, project instructions, or a slash command) to call a specific MCP tool under specific conditions, and it's up to Claude's own judgment in the moment to follow that instruction — there's no deterministic dispatcher. - MCP server registration (this part is concrete/configurable):
- Project-scoped server lives in
.mcp.jsonat the repo root, checked into git and shared with the team. - CLI:
claude mcp add --transport stdio <name> --scope project -- <command> <args...>— the--separator is required before the server command so Claude Code doesn't try to parse the server's own flags. - Example directly confirmed in the docs for a Windows/WSL stdio server:
{ "mcpServers": { "qwen-delegate": { "command": "wsl.exe", "args": ["bash", "-c", "cd ${CLAUDE_PROJECT_DIR} && ./qwen-mcp-server.sh"], "env": { "SOME_VAR": "${SOME_VAR:-default}" } } } } - Scopes, in precedence order: local (
~/.claude.json, current project only, not shared) > project (.mcp.json, shared via git) > user (~/.claude.json, all projects) > plugin-provided > claude.ai connectors. ${CLAUDE_PROJECT_DIR},${CLAUDE_PLUGIN_ROOT},${CLAUDE_PLUGIN_DATA}are available as special expansion variables in.mcp.json; general env vars use${VAR}/${VAR:-default}.- Project-scoped servers require one-time approval in interactive sessions (
claude mcp reset-project-choicesto reset;enabledMcpjsonServers/disabledMcpjsonServersin settings to pre-approve/block); non-interactive (claude -p) sessions load them without prompting unless--strict-mcp-configis passed.
- Project-scoped server lives in
Verdict: viable and directly supported — a stdio MCP server whose command is wsl.exe is an explicitly documented pattern, not a workaround.
Recommended concrete plan
-
Confirm qwen's real invocation locally first (see Open Questions — this research could not run WSL commands).
-
Write a minimal custom MCP server (Python SDK, since the ecosystem here already leans Python) with a single tool
qwen_delegate(prompt: str, timeout_s: int = 60) -> strthat:- Spawns
wsl.exe qwen -p "<prompt>"(or wraps inbash -lcif PATH/profile setup is needed forqwento resolve inside WSL). - Captures and returns stdout; on failure, returns exit code + stderr as an MCP error/tool-result so Claude Code can see it and fall back to doing the work itself.
- Enforces a timeout, since a SearXNG-backed research call could hang.
- Spawns
-
Register it as a project-scoped stdio server in
.mcp.json:{ "mcpServers": { "qwen-delegate": { "command": "wsl.exe", "args": ["python3", "/path/inside/wsl/to/qwen_mcp_server.py"] } } }or, if you keep the MCP server itself on the Windows side and only shell out to WSL for the actual qwen call, register it as a normal local Python/Node stdio command instead (
command: python,args: ["qwen_mcp_server.py"]) and have that script callwsl.exe qwen -p ...internally — this is likely simpler to debug than nesting the whole MCP server inside WSL. -
Add a routing instruction to
CLAUDE.md, e.g.:For light, low-stakes work (simple lookups, quick web/research questions, boilerplate text generation, straightforward single-file edits under ~50 lines) — prefer the
qwen_delegateMCP tool over doing it yourself, to save budget. Use your own judgment for anything requiring deep codebase context, multi-file changes, or high-stakes correctness.Remember this is a prompted convention only — there is no Claude Code config that enforces it mechanically.
Open questions / things to verify locally (could not run WSL here)
- Run
wsl qwen --helpto confirm the exact non-interactive flag name and behavior (-pvs--prompt, whether it accepts stdin piping, whether output is plain text or includes ANSI/formatting that needs stripping). - Confirm whether
qwenon this machine is literallyQwenLM/qwen-code, or Qwen-Agent, or an Ollama-served model with a custom wrapper script — runwhich qwen/ check its shebang or install location inside WSL, since this determines whether themodelProviders/settings.jsonlocal-endpoint config from the qwen-code docs actually applies. - If it is qwen-code, check
~/.qwen/settings.json(or project-level.qwen/settings.json) inside WSL for its currentmodelProviders/contextWindowSizeconfig, to confirm the ~127k context and SearXNG research module are wired up the way you expect, and to get the exact provider name to reference if you ever want to call it differently. - Decide whether the MCP server process itself should live on the Windows side (simpler
wsl.exesubprocess call, easier to iterate with Windows tooling) or inside WSL (avoids a second interop hop but complicates the.mcp.jsoncommand line) — confirm which one starts up reliably by testingclaude mcp add ... -- wsl.exe ...per the doc's example and checkingclaude mcp list//mcpinside a session. - Verify
wsl.exeis on PATH for the account Claude Code runs under, and that no interactive WSL prompt (first-run distro setup, sudo password, etc.) will block a non-interactive MCP tool call.