No longer needed - every client is a coding CLI behind the OmniRoute
gateway, not a chat UI. Drops the open-webui and qdrant services,
WEBUI_PORT/OPENWEBUI_OMNIROUTE_KEY env vars, and the openwebui-data/
qdrant-data volumes. Qdrant only ever served Open WebUI's own built-in
memory/RAG (unrelated to the gateway-level knowledgebase removed in
472e3a4), so it goes too rather than sit unused.
Docs updated: README, docs/network-access.md (ai.home/ai.haylan.ch
section was entirely about Open WebUI, rewritten around the gateway),
docs/proxy-key-onboarding.md, docs/proxy-request-priority.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VPZ6TogJiYxG8E4EQBB197
2.2 KiB
Network access: proxy.ai.home / proxy.ai.haylan.ch
This stack has no chat UI — every client is a coding CLI reaching the AI gateway (OmniRoute). It doesn't run its own reverse proxy — it publishes the gateway's API port to the host and relies on the existing Nginx Proxy Manager (NPM) instance already fronting other self-hosted services on this network.
llama.cpp's raw API stays LAN-only — deliberately
The inference API (port ${LLAMA_PORT:-8080}) is not registered in NPM and is not reachable externally. It has no authentication of its own — putting it on the public internet would mean an unauthenticated inference endpoint. Coding-agent CLIs (Claude Code, Kimi, OpenCode — see docs/coding-cli-setup.md) don't reach it directly at all now; they go through the gateway below, same as everything else.
If you later want external CLI access too, that's a deliberate scope change — see the map (issue #1) before doing it, since it changes the security posture.
The AI gateway (OmniRoute) — proxy.ai.home / proxy.ai.haylan.ch
As of issue #31 (migrated from LiteLLM), the gateway is OmniRoute:
proxy.ai.homeandproxy.ai.haylan.chboth point only at${OMNIROUTE_API_PORT:-20129}— the API port. Set up as two NPM Proxy Hosts pointing at this machine's LAN IP on that port;ai.homeinternal-only,ai.haylan.chexternal via the DMZ already forwarding to NPM (let NPM issue/manage the TLS cert as usual).- The dashboard (
${OMNIROUTE_DASHBOARD_PORT:-20128}) is never registered in NPM at all, anddocker-compose.ymlnever publishes that port to the host either — it manages every workload's keys, so it doesn't belong on the public internet, same reasoning as LiteLLM's old/ui. Unlike LiteLLM, OmniRoute's split-port mode means this is structural (no network route exists) rather than an NPM path-deny rule that has to be maintained and could be misconfigured. Reach the dashboard only from the host itself or over SSH port-forward.
Every gateway call already requires a valid API key (Bearer token, see docs/proxy-key-onboarding.md), so no extra NPM-level auth is needed for the external hostname.