update.sh now creates .env from .env.example if missing, idempotently fills in every random secret (same logic generate-secrets.sh had, now removed), resolves SEARXNG_LAN_IP from search.home via the host's own DNS, and mints OPENWEBUI_LITELLM_KEY / MEMORY_RETRIEVAL_EMBEDDING_KEY through LiteLLM's own /key/generate API once litellm is up — no more manual Admin UI step for the stack's own two workload keys. Docs updated to point at update.sh as the one command; docs/proxy-key-onboarding.md keeps the manual/API steps as the fallback and for onboarding other workloads. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
4.1 KiB
Knowledgebase, memory, and web search
Three gateway-level capabilities added on top of the AI gateway/proxy, so every client behind LiteLLM gets them — not just Open WebUI. See issue #21 for the rationale.
Not yet verified on real hardware — see issue #24.
Web search (SearXNG)
litellm-config.yaml's search_tools block wires the LAN's SearXNG instance in as a standalone REST endpoint, not a model-callable tool — call it directly:
curl http://<proxy>:4000/v1/search/searxng-search \
-H "Authorization: Bearer <a virtual key>" \
-H "Content-Type: application/json" \
-d '{"query": "...", "max_results": 5}'
Because this doesn't ask the model to emit a tool call, it sidesteps Qwen3.8-27B's known-flaky tool-calling (docs/research/qwen3.8-27b-tool-calling.md) entirely. Open WebUI's own web-search setting can point at this endpoint the same way.
Requires SEARXNG_LAN_IP set in .env so the litellm container can resolve search.home via extra_hosts — ./scripts/update.sh resolves and fills this in automatically from the host's own DNS if it's blank (use a static DHCP reservation for search.home so it doesn't drift). Full research: docs/research/litellm-searxng-search.md.
Knowledgebase (vector store / RAG)
LiteLLM's native knowledgebase feature has no Qdrant backend — the qdrant service in this stack only serves Open WebUI's own separate RAG/Memory feature and is unrelated to this — and no langchain_postgres backend either. Instead, a small in-repo service wraps LangChain's PGVector directly against its own Postgres+pgvector database (pgvector-db). This replaced an earlier attempt to vendor the third-party litellm-pgvector connector — see docs/research/langchain-pgvector-vs-litellm-pgvector.md for why. One consequence: the OpenAI-style file_search tool call on a /chat/completions request doesn't work here (no vector_store_registry entry) — query memory-retrieval directly instead (see below).
New pieces:
embedding-server— a second llama.cpp instance (small footprint,nomic-embed-text-v1.5) serving/v1/embeddings. The chat model isn't embedding-trained and llama.cpp serves one model per process, so this can't just be a flag onllama-server.pgvector-db— Postgres with the pgvector extension, separate fromlitellm-db.memory-retrieval(services/memory-retrieval/) — a ~90-line FastAPI app wrappinglangchain_postgres.PGVector, exposingPOST /ingestandPOST /query. Calls back intolitellmfor embeddings, same gateway boundary as everything else here.litellm-config.yaml'slocal-embeddingmodel entry, whichmemory-retrievalcalls through.
First-time setup
docker compose --profile tools run --rm downloader-embedding # fetch the embedding model
docker compose up -d embedding-server pgvector-db memory-retrieval
./scripts/update.sh mints MEMORY_RETRIEVAL_EMBEDDING_KEY automatically (a memory-retrieval virtual key via LiteLLM's own API) if it's blank — it calls back into litellm for embeddings, same as any other workload. See docs/proxy-key-onboarding.md if a mint fails and it needs doing by hand.
Loading memory into it
data/memory.md and data/claude-legacy-memory.md — Claude-memory-style fact files — get loaded via:
./scripts/ingest-memory.sh
One chunk per fact/paragraph line, tagged with source/section metadata. Re-run after editing either file (see the script's header comment for the no-dedup caveat).
Querying it
Directly against memory-retrieval — no in-band file_search tool call (see above):
curl -X POST http://memory-retrieval:8000/query \
-H "Authorization: Bearer <MEMORY_RETRIEVAL_API_KEY>" \
-H "Content-Type: application/json" \
-d '{"query": "...", "k": 5}'
A caller wanting RAG-augmented chat does the "search then send" pattern: query /query, splice the results into the prompt, then call litellm as normal.