haylanandClaude-Bot 52a92f6508 fix: llama-server-fast context-size exhaustion breaking Auto Mode classifier
Real failure: "Auto Mode couldn't classify this action (Classifier
stage 1 unavailable)". Reproduced directly against the server:

  {"error":{"message":"[400]: request (6186 tokens) exceeds the
  available context size (4096 tokens)"...

LLAMA_FAST_CTX_SIZE=8192 is the TOTAL across every LLAMA_FAST_PARALLEL
slot, not per-request — the main model's own .env.example comment
already calls this out, missed it when llama-server-fast was set up
(#44). With PARALLEL=2 that's 4096/slot, too small for a real
classifier call (hints + environment + recent tool-call history).

Fixed by dropping to a single slot (LLAMA_FAST_PARALLEL=1) rather than
raising ctx-size — this service doesn't need concurrent classifier
calls the way the main model needs concurrent chat sessions, so this
costs no extra VRAM. The full 8192 now goes to the one slot.

docker compose config -q validated.

Refs #5

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MrnMEdzeQzqZE5soVEXPCx
2026-09-06 22:06:59 +02:00

LLM-Server

Local AI inference stack: llama.cpp (ROCm) serving Qwen3.8-27B on an AMD Radeon AI PRO R9700, fronted by the OmniRoute AI gateway, with Lazytainer auto-suspending the inference container when idle.

See the wayfinder map (issue #1) for the full architecture rationale and open questions.

Quickstart

./scripts/update.sh

update.sh creates .env from .env.example if missing, fills in every random secret it can generate itself (via openssl, SEARXNG_LAN_IP resolved from search.home on this host), downloads the model GGUF into the models volume if it's not there yet, then pulls/builds/brings up the whole stack. Safe to re-run any time — it only fills in what's still blank, skips the model if already downloaded, and only recreates what changed.

llama.cpp's own API is internal-only — everything routes through the AI gateway below.

Pointing Claude Code CLI, Kimi CLI, or OpenCode CLI at the local endpoint: see docs/coding-cli-setup.md.

Known risk: Qwen3.8-27B's tool-calling reliability against llama.cpp's Anthropic shim is not yet verified (open upstream parser bugs against its model lineage) — see docs/research/qwen3.8-27b-tool-calling.md.

AI gateway (OmniRoute)

An AI gateway/proxy fronts llama.cpp: per-workload API keys and usage tracking. As of issue #31 this is OmniRoute, replacing the original LiteLLM setup. ./scripts/update.sh handles most of OmniRoute's secrets (see .env.example); per-workload API keys still need minting by hand in the dashboard — see docs/proxy-key-onboarding.md.

  • Gateway API: http://<this-machine>:${OMNIROUTE_PORT:-4000}/v1 locally, or proxy-ai.home / proxy-ai.haylan.ch once routed through NPM — see docs/network-access.md.
  • Dashboard (key/provider management): LAN/host-only, never published to the internet — see docs/network-access.md.
  • Issuing a key for a new workload: docs/proxy-key-onboarding.md.

Coding CLIs (see docs/coding-cli-setup.md) route through the gateway — llama-server has no published host port. Not yet verified: none of this has been smoke-tested on real hardware yet — see issue #31's tickets for the open items (provider registration, per-workload key minting).

Note on this choice: OmniRoute's own docs (docs/security/STEALTH_GUIDE.md, MITM-TPROXY-DECRYPT.md, PUBLIC_CREDS.md in its repo) describe shipped features for evading AI-provider client detection, system-wide HTTPS interception via a locally-installed root CA, and hiding credentials from secret scanners. None of that is used by this stack's configuration, but it's a real characteristic of the upstream project — see issue #31's Notes for the full research trail before extending this integration further.

The gateway also fronts SearXNG-backed web search — see docs/research/litellm-searxng-search.md for the original research (still applicable — same standalone-endpoint pattern, see issue #31's #35).

What's not here

  • Open WebUI — this stack has no chat UI; every client is a coding CLI. Removed rather than kept idle.
  • Gateway-level knowledgebase/memory (litellm-pgvector, pgvector-db, a dedicated embedding model) — removed as unwanted, unrelated to OmniRoute's own lack of parity with it (see issue #31's #34). Superseded by OmniRoute's own built-in memory feature (opt-in via the dashboard, Settings → Memory): vector store is its bundled sqlite-vec, embeddings are a local ONNX model (Transformers.js, ~400MB, downloaded into the omniroute-data volume on first use) — no external services, no static config here.
S
Description
No description provided
Readme
4.9 MiB
Languages
Shell 100%