Fix docs: no localhost, wire up ai.home / ai.haylan.ch via existing NPM
Resolves wayfinder ticket #8. The stack is reached over LAN/WAN hostnames, not localhost: - Open WebUI: ai.home (LAN) and ai.haylan.ch (external, via DMZ) routed through the existing Nginx Proxy Manager instance to the published WEBUI_PORT — no new reverse-proxy service needed in this repo. - llama.cpp's raw API stays LAN-only and unregistered in NPM (no auth of its own); coding CLIs reach it via the box's LAN IP or ai.home directly. New docs/network-access.md documents the NPM proxy-host setup and the LAN-only decision. coding-cli-setup.md and README updated to stop referencing localhost. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
@@ -2,11 +2,13 @@
|
||||
|
||||
This stack's llama.cpp server exposes two endpoints once `docker compose up` is running (see `docker-compose.yml`):
|
||||
|
||||
- **OpenAI-compatible**: `http://localhost:8080/v1` (or `${LLAMA_PORT}` if you changed it in `.env`)
|
||||
- **Anthropic Messages API shim**: `http://localhost:8080` (adds `/v1/messages`)
|
||||
- **OpenAI-compatible**: `http://<ai-box>:8080/v1` (or `${LLAMA_PORT}` if you changed it in `.env`)
|
||||
- **Anthropic Messages API shim**: `http://<ai-box>:8080` (adds `/v1/messages`)
|
||||
|
||||
Both serve the same model — `Qwen3.8-27B-UD-Q4_K_XL.gguf` — behind whichever wire format the client speaks.
|
||||
|
||||
`<ai-box>` is this machine's LAN address — its LAN IP, or `ai.home` if your local DNS resolves that hostname directly to the box. **This API is LAN-only, not reachable via `ai.haylan.ch`** — it's deliberately not registered in Nginx Proxy Manager (no auth of its own, unlike Open WebUI). See `docs/network-access.md`. If you're running a coding CLI from this machine itself, `localhost` works too.
|
||||
|
||||
> **Read this before relying on it for real work.** Qwen3.8-27B's tool-calling has **documented, open llama.cpp upstream bugs** (parser fails on text before `<tool_call>`, tool calls emitted as inert XML inside thinking blocks — see `docs/research/qwen3.8-27b-tool-calling.md`). Every setup below inherits this risk identically, regardless of which CLI or wire format you use. Don't trust it for unattended multi-step agentic work until you've run the smoke test in [issue #5](https://git.arthurerlich.de/haylan/LLM-Server/issues/5).
|
||||
|
||||
## Claude Code CLI
|
||||
@@ -14,7 +16,7 @@ Both serve the same model — `Qwen3.8-27B-UD-Q4_K_XL.gguf` — behind whichever
|
||||
Claude Code speaks the **Anthropic Messages API** — point it at the shim, not the OpenAI-compatible endpoint:
|
||||
|
||||
```bash
|
||||
export ANTHROPIC_BASE_URL=http://localhost:8080
|
||||
export ANTHROPIC_BASE_URL=http://<ai-box>:8080
|
||||
export ANTHROPIC_API_KEY=local # value is unchecked by llama.cpp, but the client requires it set
|
||||
claude
|
||||
```
|
||||
@@ -28,7 +30,7 @@ Kimi CLI speaks plain **OpenAI Chat Completions** — no shim needed. Configure
|
||||
```toml
|
||||
[providers.openai]
|
||||
type = "openai"
|
||||
base_url = "http://localhost:8080/v1"
|
||||
base_url = "http://<ai-box>:8080/v1"
|
||||
api_key = "local"
|
||||
```
|
||||
|
||||
@@ -53,7 +55,7 @@ curl -fsSL https://opencode.ai/install | bash
|
||||
"npm": "@ai-sdk/openai-compatible",
|
||||
"name": "llama.cpp (local)",
|
||||
"options": {
|
||||
"baseURL": "http://localhost:8080/v1",
|
||||
"baseURL": "http://<ai-box>:8080/v1",
|
||||
"apiKey": "sk-local-not-checked"
|
||||
},
|
||||
"models": {
|
||||
@@ -80,8 +82,8 @@ Select the model with `llamacpp/qwen3.8-27b`.
|
||||
|
||||
| CLI | Wire format | Endpoint | Config |
|
||||
|---|---|---|---|
|
||||
| Claude Code | Anthropic Messages | `http://localhost:8080` | `ANTHROPIC_BASE_URL` env var |
|
||||
| Kimi CLI | OpenAI Chat Completions | `http://localhost:8080/v1` | `config.toml` provider block |
|
||||
| OpenCode | OpenAI Chat Completions | `http://localhost:8080/v1` | `opencode.json` provider block |
|
||||
| Claude Code | Anthropic Messages | `http://<ai-box>:8080` | `ANTHROPIC_BASE_URL` env var |
|
||||
| Kimi CLI | OpenAI Chat Completions | `http://<ai-box>:8080/v1` | `config.toml` provider block |
|
||||
| OpenCode | OpenAI Chat Completions | `http://<ai-box>:8080/v1` | `opencode.json` provider block |
|
||||
|
||||
Further reading: `docs/research/qwen3.8-27b-tool-calling.md`, `docs/research/opencode-cli-setup.md`.
|
||||
|
||||
Reference in New Issue
Block a user