Not omniroute's own internal port (API_PORT stays at its default 20129, unreconfigured) - just the Docker port mapping, so existing NPM/firewall config pointed at :4000 keeps working without changes on that end. New OMNIROUTE_PORT env var is the host side of "OMNIROUTE_PORT:API_PORT" in docker-compose.yml's ports: entry. Also corrected docs/proxy-key-onboarding.md's dashboard-access instructions - DASHBOARD_PORT was never published to the host in the first place, so "http://<host>:20128" was never actually reachable as written; documented reaching it via the container's own bridge-network IP or an SSH port-forward instead. llama-server remains unexposed (no ports: entry, only expose:) - unaffected by this change, confirming it stays that way. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VPZ6TogJiYxG8E4EQBB197
19 lines
2.2 KiB
Markdown
19 lines
2.2 KiB
Markdown
# Network access: proxy.ai.home / proxy.ai.haylan.ch
|
|
|
|
This stack has no chat UI — every client is a coding CLI reaching the AI gateway (OmniRoute). It doesn't run its own reverse proxy — it publishes the gateway's API port to the host and relies on the **existing Nginx Proxy Manager (NPM)** instance already fronting other self-hosted services on this network.
|
|
|
|
## llama.cpp's raw API stays LAN-only — deliberately
|
|
|
|
The inference API (port `${LLAMA_PORT:-8080}`) is **not** registered in NPM and is **not** reachable externally. It has no authentication of its own — putting it on the public internet would mean an unauthenticated inference endpoint. Coding-agent CLIs (Claude Code, Kimi, OpenCode — see `docs/coding-cli-setup.md`) don't reach it directly at all now; they go through the gateway below, same as everything else.
|
|
|
|
If you later want external CLI access too, that's a deliberate scope change — see the map ([issue #1](https://git.arthurerlich.de/haylan/LLM-Server/issues/1)) before doing it, since it changes the security posture.
|
|
|
|
## The AI gateway (OmniRoute) — `proxy.ai.home` / `proxy.ai.haylan.ch`
|
|
|
|
As of [issue #31](https://git.arthurerlich.de/haylan/LLM-Server/issues/31) (migrated from LiteLLM), the gateway is OmniRoute:
|
|
|
|
- **`proxy.ai.home`** and **`proxy.ai.haylan.ch`** both point only at `${OMNIROUTE_PORT:-4000}` — the API port. Set up as two NPM Proxy Hosts pointing at this machine's LAN IP on that port; `ai.home` internal-only, `ai.haylan.ch` external via the DMZ already forwarding to NPM (let NPM issue/manage the TLS cert as usual).
|
|
- The **dashboard** (`${OMNIROUTE_DASHBOARD_PORT:-20128}`) is never registered in NPM at all, and `docker-compose.yml` never publishes that port to the host either — it manages every workload's keys, so it doesn't belong on the public internet, same reasoning as LiteLLM's old `/ui`. Unlike LiteLLM, OmniRoute's split-port mode means this is structural (no network route exists) rather than an NPM path-deny rule that has to be maintained and could be misconfigured. Reach the dashboard only from the host itself or over SSH port-forward.
|
|
|
|
**Every gateway call already requires a valid API key** (Bearer token, see `docs/proxy-key-onboarding.md`), so no extra NPM-level auth is needed for the external hostname.
|