haylanandClaude-Bot 828bd4c046 feat(llm): dedicate a CPU-only backend for the qwen-code tool-call classifier
fastModel in ~/.qwen/settings.json (permissions.autoMode.classifier) was
aliased onto llama-server's own 27B connection, so every tool-call safety
check queued behind whatever heavy generation was already running on that
model's 2 GPU slots.

Add qwen-classifier: a separate llama.cpp instance, CPU-only, running
Qwen3-4B-Instruct-2507 (the smallest Qwen3 with native >=131072 context,
qwen-code's requirement, without lossy RoPE scaling). Structurally isolated
from llama-server's queue instead of sharing it. Sized for gameserver's
~17GiB free system RAM: q8_0/q8_0 KV at full 131072 ctx (~9.8GiB) + Q4_K_M-
class weights (~2.3GiB) fits comfortably, with better KV quality than the
q4_0 that would've been needed to fit this on the GPU's ~6GiB free VRAM.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-09 16:26:51 +02:00

LLM-Server

Local AI inference stack: llama.cpp (ROCm) serving Qwen3.8-27B on an AMD Radeon AI PRO R9700, fronted by the OmniRoute AI gateway, with Lazytainer auto-suspending the inference container when idle.

See the wayfinder map (issue #1) for the full architecture rationale and open questions.

Quickstart

./scripts/update.sh

update.sh creates .env from .env.example if missing, fills in every random secret it can generate itself (via openssl, SEARXNG_LAN_IP resolved from search.home on this host), downloads the model GGUF into the models volume if it's not there yet, then pulls/builds/brings up the whole stack. Safe to re-run any time — it only fills in what's still blank, skips the model if already downloaded, and only recreates what changed.

llama.cpp's own API is internal-only — everything routes through the AI gateway below.

Pointing Claude Code CLI, Kimi CLI, OpenCode CLI, or Qwen Code CLI at the local endpoint: see docs/coding-cli-setup/.

Known risk: Qwen3.8-27B's tool-calling reliability against llama.cpp's Anthropic shim is not yet verified (open upstream parser bugs against its model lineage) — see docs/research/qwen3.8-27b-tool-calling.md.

AI gateway (OmniRoute)

An AI gateway/proxy fronts llama.cpp: per-workload API keys and usage tracking. As of issue #31 this is OmniRoute, replacing the original LiteLLM setup. ./scripts/update.sh handles most of OmniRoute's secrets (see .env.example); per-workload API keys still need minting by hand in the dashboard — see docs/proxy-key-onboarding.md.

  • Gateway API: http://<this-machine>:${OMNIROUTE_PORT:-4000}/v1 locally, or proxy-ai.home / proxy-ai.haylan.ch once routed through NPM — see docs/network-access.md.
  • Dashboard (key/provider management): LAN/host-only, never published to the internet — see docs/network-access.md.
  • Issuing a key for a new workload: docs/proxy-key-onboarding.md.

Coding CLIs (see docs/coding-cli-setup/) route through the gateway — llama-server has no published host port. Not yet verified: none of this has been smoke-tested on real hardware yet — see issue #31's tickets for the open items (provider registration, per-workload key minting).

Note on this choice: OmniRoute's own docs (docs/security/STEALTH_GUIDE.md, MITM-TPROXY-DECRYPT.md, PUBLIC_CREDS.md in its repo) describe shipped features for evading AI-provider client detection, system-wide HTTPS interception via a locally-installed root CA, and hiding credentials from secret scanners. None of that is used by this stack's configuration, but it's a real characteristic of the upstream project — see issue #31's Notes for the full research trail before extending this integration further.

The gateway also fronts SearXNG-backed web search — see docs/research/litellm-searxng-search.md for the original research (still applicable — same standalone-endpoint pattern, see issue #31's #35).

S
Description
No description provided
Readme
4.9 MiB
Languages
Shell 100%