Files
LLM-Server/docs/research/qwen3.8-27b-tool-calling.md

14 KiB
Raw Permalink Blame History

Is Qwen3.8-27B's tool-call output compatible with llama.cpp's Anthropic Messages API shim?

Date: 2026-08-24 Answer: No — not out of the box, with meaningful confidence. Evidence points to the same class of problem Qwen3-Coder had (issue #176 / unsloth GGUF discussion #10), but the failure this time sits one layer lower: llama.cpp's own tool-call grammar/parser for the Qwen3.5 architecture family (which Qwen3.8 is built on) is documented as broken for agentic tool-calling as of the latest issues found, independent of which client (Anthropic shim vs. OpenAI-compatible endpoint) is used. This is worse than "needs a chat-template swap" — it's "upstream parser doesn't reliably extract the tool call at all in some conditions."

1. Qwen3.8-27B model card (HF: Qwen/Qwen3.8-27B)

  • The model card does not publish a dedicated tool-call syntax/spec, nor call out a specific chat_template.jinja for tool use in the text pulled. It describes agentic improvements only in generic marketing language ("stronger autonomous planning and better handling of environment feedback"), not a documented wire format. Source: https://huggingface.co/Qwen/Qwen3.8-27B
  • Architecture: dense (not MoE), hybrid linear-attention design — "16 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))" per one summary of the card, i.e. mostly Gated DeltaNet (linear attention) layers with a minority of standard attention layers. It is explicitly built "on the architectural foundation of Qwen3.5" (per search-indexed card text) and is a native VLM (images + video), not a coding-specialist MoE like Qwen3-Coder. Source: https://huggingface.co/Qwen/Qwen3.8-27B ; corroborated by https://github.com/ggml-org/llama.cpp/discussions/27164 (48 Gated DeltaNet layers + 16 full-attention layers across 64 total layers).
  • Confidence: medium. The model card content here was retrieved via an automated fetch/summarize pass, not read verbatim by a human — the architecture string and "foundation of Qwen3.5" phrasing should be treated as paraphrase-level accurate, not a verbatim quote guarantee. No independent second source confirms the exact chat-template filename question either way (i.e., I could not confirm whether a chat_template.jinja tool-call section exists or not — the fetch reported none found, but that is a negative result from a lossy summarization step, not a certainty).

2. llama.cpp's Anthropic Messages API shim (PR #17570)

  • The shim adds /v1/messages and /v1/messages/count_tokens, and works by converting Anthropic's wire format to llama.cpp's existing OpenAI-compatible internal format, then reusing the existing inference/tool-calling pipeline — it does not implement its own separate tool-call extraction logic. Source: https://github.com/ggml-org/llama.cpp/pull/17570
  • That means the shim's tool-call reliability is entirely inherited from llama.cpp's normal chat completions tool-calling path (chat template + jinja + grammar-constrained parsing), not something Anthropic-shim-specific. Whatever breaks tool-calling on /v1/chat/completions for a given model breaks it identically through /v1/messages.
  • llama.cpp's tool-calling is model-agnostic by design in principle: it uses --jinja to process the model's own chat template, and falls back to a "generic tool call style" when a model has no native/recognized format. In principle this makes it work with "pretty much any model" (per the PR author's own comment). But "model-agnostic in design" is not the same as "bug-free for every model" — see section 3. Source: https://github.com/ggml-org/llama.cpp/pull/17570 ; https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
  • The PR author's own tested-models list at the time was Qwen3 Coder 30B-A3B and gpt-oss-20b ("decent results for simple tasks"); Qwen3.8/Qwen3.5 is not mentioned as tested in the PR itself. Source: https://github.com/ggml-org/llama.cpp/pull/17570

3. Qwen3.5-lineage tool-calling bugs in llama.cpp (direct evidence — this is the crux)

Qwen3.8-27B explicitly shares its architecture lineage with Qwen3.5. Multiple open or only partially-fixed llama.cpp issues document broken tool-call parsing specifically for that lineage:

  • Issue #21158 — "[Bug]: Qwen3.5-27B tool call parsing still broken after PR #20424 (build 8576)." Symptom: tool-call parsing fails whenever the model emits any text before the <tool_call> XML tag ("Failed to parse input at pos 41: <tool_call>..."). PR #20424 landed as a partial fix (proper reasoning-tag prefill reading) but did not fully resolve it; the issue was closed by being folded into a broader still-open issue (#20260). Reporter states this "makes Qwen3.5 unusable for agentic tool-calling workflows with llama-server," forcing users to disable structured tools entirely. Source: https://github.com/ggml-org/llama.cpp/issues/21158
  • Issue #20837 — "Qwen3.5 9B often prints tool calls in XML and stops when thinking is enabled - tool calls inside thinking block." Status: open, no fix/PR attached at time of research. Symptom: with thinking enabled, the model's tool call gets emitted as literal XML text inside/around the reasoning block and the turn ends without the call being recognized/executed; workaround is enable_thinking:false. Present since llama.cpp build 8255. Source: https://github.com/ggml-org/llama.cpp/issues/20837
  • Related, same architecture family and directly analogous failure mode to what Qwen3-Coder hit: Issue #20164 (tool calling repeatedly fails under long context when a tool has multiple optional parameters — affects Qwen3.5-35B and Qwen3-Coder-Next) and Issue #20809 (Qwen3-Instruct-2507 tool calls misclassified as reasoning_content due to false "thinking" detection). Sources: https://github.com/ggml-org/llama.cpp/issues/20164 , https://github.com/ggml-org/llama.cpp/issues/20809

None of these issues are Qwen3.8-27B-specific by name (Qwen3.8-27B is very new — released ~Aug 13-14, 2026 per search results), but they are architecture-lineage bugs: the failure is in llama.cpp's XML/<tool_call> grammar-trigger and thinking-block parsing logic for the Qwen3.5-style chat template and tool-call format, which Qwen3.8 explicitly inherits ("built on the architectural foundation of Qwen3.5"). There is no evidence found that Qwen3.8-27B introduces a different tool-call wire format that would sidestep these bugs.

4. Comparison to the Qwen3-Coder precedent (issue #176 / unsloth discussion #10)

  • qwen-code issue #176: qwen3-coder-30b-a3b via llama.cpp emitted syntactically-correct XML tool-call markup, but the client-side parsing pipeline reported tool_calls_count=0 — the call was silently dropped. No fix is recorded in the issue itself. Source: https://github.com/QwenLM/qwen-code/issues/176
  • unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF discussion #10 ("New Chat Template + Tool Calling Fixes as of 05 Aug, 2025"): Unsloth identified this as a universal bug affecting all GGUF uploads of Qwen3-Coder regardless of who quantized them, not an unsloth-specific packaging error, and shipped a fixed chat template usable via --jinja / --chat-template-file, after coordinating with the Qwen team. Source: https://huggingface.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF/discussions/10
  • Contrast for Qwen3.8/Qwen3.5: The Qwen3-Coder fix was a chat-template-level fix (swap in a corrected Jinja template) that fully resolved the issue once applied. For the Qwen3.5 lineage, issue #21158 shows a partial upstream PR (#20424) landed and still did not fully fix it — the bug was reclassified into a broader still-open parser issue. This suggests the Qwen3.5/3.8 problem may not be fixable purely by dropping in a corrected chat template the way Qwen3-Coder's was; it may require changes to llama.cpp's grammar-trigger/parser code itself (specifically: not triggering XML tool-call parsing when it occurs inside or straddles reasoning/thinking blocks, and tolerating leading text before <tool_call>). A separate discussion also confirms upstream chat-template churn/incompatibility for this lineage: unsloth/Qwen3.5-35B-A3B-GGUF discussion #18 ("LM Studio does not support the newly updated chat template") — LM Studio's Jinja engine threw Unknown test: sequence on Unsloth's updated Qwen3.5 template and silently fell back to a default template, breaking tool-calling until a community member posted a patched template. This shows the Qwen3.5 template itself has already churned at least once post-release in ways that broke naive Jinja engines — a second, independent axis of fragility beyond the llama.cpp parser bugs above. Source: https://huggingface.co/unsloth/Qwen3.5-35B-A3B-GGUF/discussions/18

5. Does the hybrid linear-attention / VLM-dense architecture (vs. coding-MoE) have documented bearing on tool-call reliability?

No concrete causal link found — flagging as untested/unknown, with one adjacent data point:

  • github.com/ggml-org/llama.cpp/discussions/27164 documents a real Qwen3.8-27B bug on llama.cpp, but it is a CUDA numerical-correctness bug in the DeltaNet execution path (corrupted/garbage output, root-caused to a stale libggml-cuda.so / stale build rather than a stale git clone reporting falsely up-to-date), not a tool-calling/parsing bug. It's evidence that the new hybrid architecture's backend support is young and has had correctness bugs, which is relevant risk context (a model that occasionally emits garbage tokens will also occasionally emit garbage tool calls), but it is not evidence that linear attention specifically degrades tool-call formatting reliability in agentic loops. Source: https://github.com/ggml-org/llama.cpp/discussions/27164
  • No primary source (Qwen model card, llama.cpp issues, or Anthropic docs) makes a claim connecting "hybrid linear attention" or "VLM-capable dense" architecture directly to tool-call reliability in agentic loops one way or the other. The tool-call bugs found (section 3) are chat-template/parser bugs shared across the Qwen3.5 lineage (which includes non-hybrid variants too, e.g. Qwen3.5-9B, Qwen3-Coder-Next), so they track the shared chat-template/tool-call-format design, not the attention mechanism specifically. Treat any claim of an architecture→reliability causal link as unconfirmed.

Bottom line for the deployment decision

  1. Qwen3.8-27B does not have a documented, stable, llama.cpp-verified tool-call format as of this research (2026-08-24) — it's ~10 days post-release at time of writing.
  2. llama.cpp's Anthropic Messages API shim (PR #17570) is a thin format-converter over the existing OpenAI-compatible tool-calling pipeline — it adds no independent risk beyond what /v1/chat/completions already has for this model, but it also inherits every bug that pipeline has for the Qwen3.5 lineage.
  3. Concretely, expect the same class of failure mode Qwen3-Coder hit (#176): tool calls syntactically emitted by the model but not recognized by the harness — plus a lineage-specific failure mode where thinking-enabled generation emits the tool call as inert XML text inside/around the reasoning block (#20837) or fails to parse when the model prefixes any text before <tool_call> (#21158, only partially fixed by #20424).
  4. Unlike Qwen3-Coder, the fix path here is not confirmed to be a simple chat-template swap — the most specific issue (#21158) shows a merged upstream PR that didn't fully resolve it, with the residual bug folded into a broader still-open llama.cpp issue (#20260). A workaround exists (enable_thinking:false, or disabling structured tools) but that trades away thinking-mode quality or agentic capability entirely.
  5. Recommendation: keep the wayfinder map's existing "unverified, accepted knowingly" framing, but sharpen it — this is not merely "untested," there is now direct, dated evidence of upstream llama.cpp parser bugs against this exact model lineage as of the current llama.cpp master. Before relying on Qwen3.8-27B for unattended agentic tool use via the Anthropic shim, this needs a live smoke test against the actual deployed llama.cpp build (not just docs) — specifically: a multi-turn tool-calling test with thinking enabled, since that is where the two most concrete bugs live.

Sources

Confidence/uncertainty summary

  • High confidence: the Anthropic shim reuses the OpenAI-compatible tool-calling pipeline (stated directly in PR #17570); the Qwen3-Coder precedent bug and its chat-template fix (directly documented in issue #176 and discussion #10); the existence of open/partially-fixed Qwen3.5-lineage tool-call parser bugs in llama.cpp (issues #21158, #20837, directly documented, with issue numbers and quoted symptoms).
  • Medium confidence: that Qwen3.8-27B's chat template is unchanged from Qwen3.5's in ways that make it inherit these exact bugs — this is inferred from "built on the architectural foundation of Qwen3.5" language, not confirmed by diffing the actual chat_template.jinja files of Qwen3.5 vs. Qwen3.8-27B (not attempted in this pass).
  • Low confidence / unconfirmed: any claim that the hybrid linear-attention architecture itself (as opposed to the shared chat-template/parser design) causes tool-call unreliability — no primary source makes this connection; flagged as untested per the ticket's own instruction.
  • Not independently verified: the exact current-master status of issues #21158/#20260/#20837 (i.e. whether they have since been fixed after this research's fetch date) — GitHub issue states can change daily; re-check before final go/no-go if this decision is time-sensitive.