Move docs/coding-cli-setup.md to docs/coding-cli-setup/ with one file per CLI (claude-code, kimi-cli, opencode, qwen-code) plus a shared index.md for the gateway intro, tool-calling risk note, and summary table. Also fixes the qwen-code doc: context sizes are per-slot (LLAMA_CTX_SIZE / LLAMA_PARALLEL), not raw LLAMA_CTX_SIZE (same fix applied to OpenCode's limit.context); documents the fastModel classifier provider and its own context math; adds the omniroute-search MCP server (SearXNG web search) and Auto Mode permissions tuning that were missing from the original qwen-code section. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2.9 KiB
Pointing a coding-agent CLI at this stack
This stack routes through the AI gateway (OmniRoute, see issue #31) rather than talking to llama.cpp directly — llama.cpp's own port is internal-only now (see docker-compose.yml). The gateway exposes:
- OpenAI-compatible:
http://<ai-box>:${OMNIROUTE_PORT:-4000}/v1 - Anthropic Messages API (OmniRoute's own
/v1/messagesendpoint, translating to the OpenAI-compatible backend):http://<ai-box>:${OMNIROUTE_PORT:-4000}
Both serve the same underlying model — Qwen3.8-27B-UD-Q4_K_XL.gguf, registered in the gateway (naming is yours to pick when adding the llama-cpp provider connection — these docs assume qwen3.8-27b-local for continuity) — behind whichever wire format the client speaks.
<ai-box> is this machine's LAN address, or proxy-ai.home if your local DNS resolves that hostname directly to the box — see docs/network-access.md. If you're running a coding CLI from this machine itself, localhost works too.
Each CLI needs its own virtual key — create one per docs/proxy-key-onboarding.md (omniroute's dashboard, <workload>-<purpose> naming, e.g. claude-code-cli, kimi-cli, opencode-cli). No budget set by default. These are the machine's interactive/high-priority workloads per docs/proxy-request-priority.md.
Read this before relying on it for real work. Qwen3.8-27B's tool-calling has documented, open llama.cpp upstream bugs (parser fails on text before
<tool_call>, tool calls emitted as inert XML inside thinking blocks — seedocs/research/qwen3.8-27b-tool-calling.md). Every CLI below inherits this risk identically, regardless of wire format. Don't trust it for unattended multi-step agentic work until you've run the smoke test in issue #5 (and the proxy-specific smoke test in issue #17).
Per-CLI setup
Summary
| CLI | Wire format | Endpoint | Config |
|---|---|---|---|
| Claude Code | Anthropic Messages | http://<ai-box>:${OMNIROUTE_PORT:-4000} |
ANTHROPIC_BASE_URL env var |
| Kimi CLI | OpenAI Chat Completions | http://<ai-box>:${OMNIROUTE_PORT:-4000}/v1 |
config.toml provider block |
| OpenCode | OpenAI Chat Completions | http://<ai-box>:${OMNIROUTE_PORT:-4000}/v1 |
opencode.json provider block |
| Qwen Code | OpenAI Chat Completions (2 models: chat + fastModel) |
http://<ai-box>:${OMNIROUTE_PORT:-4000}/v1 |
~/.qwen/settings.json modelProviders.openai |
Further reading: docs/research/qwen3.8-27b-tool-calling.md, docs/proxy-key-onboarding.md.