Migrate Open WebUI and coding CLIs to the AI proxy (resolves #15)

Open WebUI now points at litellm instead of llama-server directly, using a
provisioned virtual key. llama-server's host port is dropped (internal-only
on the ai-stack network) since the proxy is the only intended entry point
now. docs/coding-cli-setup.md repointed at the proxy's endpoints/ports with
per-CLI virtual keys instead of the old shared dummy key.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-25 07:18:03 +02:00
co-authored by Claude-Bot
parent 0aefb36a48
commit 5cb34b19f3
5 changed files with 52 additions and 36 deletions
+24 -22
View File
@@ -1,23 +1,25 @@
# Pointing a coding-agent CLI at this stack
This stack's llama.cpp server exposes two endpoints once `docker compose up` is running (see `docker-compose.yml`):
This stack routes through the [AI proxy](https://git.arthurerlich.de/haylan/LLM-Server/issues/9) (LiteLLM) rather than talking to llama.cpp directly — llama.cpp's own port is internal-only now (see `docker-compose.yml`). The proxy exposes:
- **OpenAI-compatible**: `http://<ai-box>:8080/v1` (or `${LLAMA_PORT}` if you changed it in `.env`)
- **Anthropic Messages API shim**: `http://<ai-box>:8080` (adds `/v1/messages`)
- **OpenAI-compatible**: `http://<ai-box>:4000/v1` (or `${LITELLM_PORT}` if you changed it in `.env`)
- **Anthropic Messages API** (LiteLLM's own unified `/v1/messages` endpoint, translating to the OpenAI-compatible backend): `http://<ai-box>:4000`
Both serve the same model — `Qwen3.8-27B-UD-Q4_K_XL.gguf` — behind whichever wire format the client speaks.
Both serve the same underlying model — `Qwen3.8-27B-UD-Q4_K_XL.gguf`, registered in the proxy as `qwen3.8-27b-local` — behind whichever wire format the client speaks.
`<ai-box>` is this machine's LAN address — its LAN IP, or `ai.home` if your local DNS resolves that hostname directly to the box. **This API is LAN-only, not reachable via `ai.haylan.ch`** — it's deliberately not registered in Nginx Proxy Manager (no auth of its own, unlike Open WebUI). See `docs/network-access.md`. If you're running a coding CLI from this machine itself, `localhost` works too.
`<ai-box>` is this machine's LAN address, or `proxy.ai.home` if your local DNS resolves that hostname directly to the box — see `docs/network-access.md`. If you're running a coding CLI from this machine itself, `localhost` works too.
> **Read this before relying on it for real work.** Qwen3.8-27B's tool-calling has **documented, open llama.cpp upstream bugs** (parser fails on text before `<tool_call>`, tool calls emitted as inert XML inside thinking blocks — see `docs/research/qwen3.8-27b-tool-calling.md`). Every setup below inherits this risk identically, regardless of which CLI or wire format you use. Don't trust it for unattended multi-step agentic work until you've run the smoke test in [issue #5](https://git.arthurerlich.de/haylan/LLM-Server/issues/5).
**Each CLI needs its own virtual key** — create one per docs/proxy-key-onboarding.md (LiteLLM's Admin UI, `<workload>-<purpose>` naming, e.g. `claude-code-cli`, `kimi-cli`, `opencode-cli`). No budget set by default. These are the machine's interactive/high-priority workloads per `docs/proxy-request-priority.md`.
> **Read this before relying on it for real work.** Qwen3.8-27B's tool-calling has **documented, open llama.cpp upstream bugs** (parser fails on text before `<tool_call>`, tool calls emitted as inert XML inside thinking blocks — see `docs/research/qwen3.8-27b-tool-calling.md`). Every setup below inherits this risk identically, regardless of which CLI or wire format you use. Don't trust it for unattended multi-step agentic work until you've run the smoke test in [issue #5](https://git.arthurerlich.de/haylan/LLM-Server/issues/5) (and the proxy-specific smoke test in [issue #17](https://git.arthurerlich.de/haylan/LLM-Server/issues/17)).
## Claude Code CLI
Claude Code speaks the **Anthropic Messages API** — point it at the shim, not the OpenAI-compatible endpoint:
Claude Code speaks the **Anthropic Messages API** — point it at the proxy's unified endpoint, not llama.cpp directly:
```bash
export ANTHROPIC_BASE_URL=http://<ai-box>:8080
export ANTHROPIC_API_KEY=local # value is unchecked by llama.cpp, but the client requires it set
export ANTHROPIC_BASE_URL=http://<ai-box>:4000
export ANTHROPIC_API_KEY=<claude-code-cli virtual key>
claude
```
@@ -25,13 +27,13 @@ Requires llama.cpp's `--jinja` flag (already set in `docker-compose.yml`) — wi
## Kimi CLI
Kimi CLI speaks plain **OpenAI Chat Completions** — no shim needed. Configure a provider block in its config file (`config.toml`):
Kimi CLI speaks plain **OpenAI Chat Completions**. Configure a provider block in its config file (`config.toml`):
```toml
[providers.openai]
type = "openai"
base_url = "http://<ai-box>:8080/v1"
api_key = "local"
base_url = "http://<ai-box>:4000/v1"
api_key = "<kimi-cli virtual key>"
```
If Kimi CLI's response parsing gets confused by Qwen's `<think>...</think>` reasoning tags, check its `reasoning_key` setting — it's configurable for non-standard local server responses.
@@ -51,15 +53,15 @@ curl -fsSL https://opencode.ai/install | bash
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"llamacpp": {
"aiproxy": {
"npm": "@ai-sdk/openai-compatible",
"name": "llama.cpp (local)",
"name": "AI proxy (local)",
"options": {
"baseURL": "http://<ai-box>:8080/v1",
"apiKey": "sk-local-not-checked"
"baseURL": "http://<ai-box>:4000/v1",
"apiKey": "<opencode-cli virtual key>"
},
"models": {
"qwen3.8-27b": {
"qwen3.8-27b-local": {
"name": "Qwen3.8-27B",
"limit": { "context": 65536, "output": 8192 }
}
@@ -71,7 +73,7 @@ curl -fsSL https://opencode.ai/install | bash
Set `limit.context` to match whatever `LLAMA_CTX_SIZE` this stack is actually running with (`.env`), not a value assumed from the model card — OpenCode uses it for its own context-management bookkeeping, not the server.
Select the model with `llamacpp/qwen3.8-27b`.
Select the model with `aiproxy/qwen3.8-27b-local`.
**OpenCode-specific risks** (on top of the shared Qwen3.8-27B tool-calling risk above):
- Requires llama.cpp's `--jinja` flag (already set) — without it, OpenCode's unconditional tool-call scaffolding gets a 500.
@@ -82,8 +84,8 @@ Select the model with `llamacpp/qwen3.8-27b`.
| CLI | Wire format | Endpoint | Config |
|---|---|---|---|
| Claude Code | Anthropic Messages | `http://<ai-box>:8080` | `ANTHROPIC_BASE_URL` env var |
| Kimi CLI | OpenAI Chat Completions | `http://<ai-box>:8080/v1` | `config.toml` provider block |
| OpenCode | OpenAI Chat Completions | `http://<ai-box>:8080/v1` | `opencode.json` provider block |
| Claude Code | Anthropic Messages | `http://<ai-box>:4000` | `ANTHROPIC_BASE_URL` env var |
| Kimi CLI | OpenAI Chat Completions | `http://<ai-box>:4000/v1` | `config.toml` provider block |
| OpenCode | OpenAI Chat Completions | `http://<ai-box>:4000/v1` | `opencode.json` provider block |
Further reading: `docs/research/qwen3.8-27b-tool-calling.md`, `docs/research/opencode-cli-setup.md`.
Further reading: `docs/research/qwen3.8-27b-tool-calling.md`, `docs/research/opencode-cli-setup.md`, `docs/proxy-key-onboarding.md`.