haylanandClaude-Bot cb9b3a9045 fix(docs): update stale MEMORY_RETRIEVAL_EMBEDDING_KEY refs after litellm-pgvector revert
README.md and docs/proxy-key-onboarding.md were edited by later commits
(b996b1f, 24d749b) that the litellm-pgvector revert didn't touch, so they
still named the now-removed MEMORY_RETRIEVAL_EMBEDDING_KEY var. Left
docs/research/langchain-pgvector-vs-litellm-pgvector.md as-is — it's a
dated record of the decision being reversed, not live config.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WHfjWrSEcGhCoeu6dQfDa
2026-09-02 22:46:10 +02:00

LLM-Server

Local AI inference stack: llama.cpp (ROCm) serving Qwen3.8-27B on an AMD Radeon AI PRO R9700, fronted by Open WebUI (RAG + Memory via Qdrant), with Lazytainer auto-suspending the inference container when idle.

See the wayfinder map (issue #1) for the full architecture rationale and open questions.

Quickstart

./scripts/download-model.sh
./scripts/update.sh

update.sh creates .env from .env.example if missing, fills in every secret and per-workload virtual key it can generate itself (random secrets via openssl, OPENWEBUI_LITELLM_KEY/LITELLM_PGVECTOR_EMBEDDING_KEY minted through LiteLLM's own /key/generate API, SEARXNG_LAN_IP resolved from search.home on this host), then pulls/builds/brings up the whole stack. Safe to re-run any time — it only fills in what's still blank and only recreates what changed. See docs/proxy-key-onboarding.md if a key mint fails and needs doing by hand.

  • Open WebUI: http://<this-machine>:3000 locally, or ai.home / ai.haylan.ch once routed through Nginx Proxy Manager — see docs/network-access.md. First signup becomes the admin account (WEBUI_AUTH is on).
  • llama.cpp's own API is internal-only now — everything routes through the AI proxy below.

Pointing Claude Code CLI, Kimi CLI, or OpenCode CLI at the local endpoint: see docs/coding-cli-setup.md.

Known risk: Qwen3.8-27B's tool-calling reliability against llama.cpp's Anthropic shim is not yet verified (open upstream parser bugs against its model lineage) — see docs/research/qwen3.8-27b-tool-calling.md.

AI proxy (LiteLLM)

An AI gateway/proxy fronts llama.cpp: per-workload virtual keys, usage tracking, and a shadow cost estimate ("what this would have cost on Claude Sonnet 5"). ./scripts/update.sh handles LITELLM_MASTER_KEY/LITELLM_SALT_KEY and every other secret (see .env.example).

Open WebUI and the coding CLIs (see docs/coding-cli-setup.md) route through the proxy now — llama-server has no published host port anymore. Not yet verified: none of this has been smoke-tested on real hardware (LiteLLM's priority scheduler in particular is beta — see docs/proxy-request-priority.md) — see issue #17.

Web search, knowledgebase, and memory

The gateway also fronts SearXNG-backed web search and a pgvector-backed knowledgebase (loaded with data/memory.md / data/claude-legacy-memory.md), wired at the LiteLLM layer so every client gets them, not just Open WebUI — see docs/memory-knowledgebase.md. Not yet verified on real hardware — see issue #24.

S
Description
No description provided
Readme
4.9 MiB
Languages
Shell 100%