haylan 213550e44b fix(litellm): prefix EMBEDDING__MODEL with openai/ for litellm-pgvector
litellm.aembedding can't infer a provider from a bare model name plus a
custom api_base — litellm-pgvector's embedding_service.py was hitting
'litellm.BadRequestError: LLM Provider NOT provided' on every query-time
embedding (i.e. every search). Same openai/ prefix already used for
qwen3.8-27b-local and local-embedding in litellm-config.yaml.
2026-09-02 21:32:23 +00:00

LLM-Server

Local AI inference stack: llama.cpp (ROCm) serving Qwen3.8-27B on an AMD Radeon AI PRO R9700, fronted by Open WebUI (RAG + Memory via Qdrant), with Lazytainer auto-suspending the inference container when idle.

See the wayfinder map (issue #1) for the full architecture rationale and open questions.

Quickstart

./scripts/update.sh

update.sh creates .env from .env.example if missing, fills in every secret and per-workload virtual key it can generate itself (random secrets via openssl, OPENWEBUI_LITELLM_KEY/LITELLM_PGVECTOR_EMBEDDING_KEY minted through LiteLLM's own /key/generate API, SEARXNG_LAN_IP resolved from search.home on this host), downloads both model GGUFs into the models volume if they're not there yet, then pulls/builds/brings up the whole stack. Safe to re-run any time — it only fills in what's still blank, skips models already downloaded, and only recreates what changed. See docs/proxy-key-onboarding.md if a key mint fails and needs doing by hand.

  • Open WebUI: http://<this-machine>:3000 locally, or ai.home / ai.haylan.ch once routed through Nginx Proxy Manager — see docs/network-access.md. First signup becomes the admin account (WEBUI_AUTH is on).
  • llama.cpp's own API is internal-only now — everything routes through the AI proxy below.

Pointing Claude Code CLI, Kimi CLI, or OpenCode CLI at the local endpoint: see docs/coding-cli-setup.md.

Known risk: Qwen3.8-27B's tool-calling reliability against llama.cpp's Anthropic shim is not yet verified (open upstream parser bugs against its model lineage) — see docs/research/qwen3.8-27b-tool-calling.md.

AI proxy (LiteLLM)

An AI gateway/proxy fronts llama.cpp: per-workload virtual keys, usage tracking, and a shadow cost estimate ("what this would have cost on Claude Sonnet 5"). ./scripts/update.sh handles LITELLM_MASTER_KEY/LITELLM_SALT_KEY and every other secret (see .env.example).

Open WebUI and the coding CLIs (see docs/coding-cli-setup.md) route through the proxy now — llama-server has no published host port anymore. Not yet verified: none of this has been smoke-tested on real hardware (LiteLLM's priority scheduler in particular is beta — see docs/proxy-request-priority.md) — see issue #17.

Web search, knowledgebase, and memory

The gateway also fronts SearXNG-backed web search and a pgvector-backed knowledgebase (loaded with data/memory.md / data/claude-legacy-memory.md), wired at the LiteLLM layer so every client gets them, not just Open WebUI — see docs/memory-knowledgebase.md. Not yet verified on real hardware — see issue #24.

S
Description
No description provided
Readme
4.9 MiB
Languages
Shell 100%