Files
LLM-Server/docs/memory-knowledgebase.md
T
haylanandClaude-Bot f508f5670a feat(litellm): wire SearXNG search, pgvector knowledgebase, and memory ingestion
- search_tools block in litellm-config.yaml (SearXNG as a first-class
  search_provider, standalone /v1/search endpoint, not a model tool) plus
  extra_hosts on the litellm service so it can resolve search.home.
- New embedding-server (nomic-embed-text-v1.5 on a second llama.cpp
  instance), pgvector-db, and litellm-pgvector services — LiteLLM's native
  knowledgebase feature has no Qdrant backend, so this is the only
  self-hosted path (docs/research/litellm-knowledgebase.md).
- vector_store_registry + local-embedding model entry in
  litellm-config.yaml, wiring it together.
- scripts/ingest-memory.sh to load data/memory.md and
  data/claude-legacy-memory.md into the knowledgebase.
- docs/memory-knowledgebase.md documenting the whole setup; data/
  gitignored (personal memory content, not meant to be committed).
- New .env vars (SEARXNG_LAN_IP, PGVECTOR_DB_PASSWORD,
  LITELLM_PGVECTOR_API_KEY, LITELLM_PGVECTOR_EMBEDDING_KEY,
  EMBEDDING_MODEL_FILE) and generate-secrets.sh support for the
  auto-generatable ones.

Resolves #22 and #23 (wayfinder map #21). Not yet verified on real
hardware — see #24.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-02 21:38:16 +02:00

3.6 KiB

Knowledgebase, memory, and web search

Three gateway-level capabilities added on top of the AI gateway/proxy, so every client behind LiteLLM gets them — not just Open WebUI. See issue #21 for the rationale.

Not yet verified on real hardware — see issue #24. In particular: litellm-pgvector's Prisma migrations on first boot, and the exact vector_store_registry field names for the pg_vector provider.

Web search (SearXNG)

litellm-config.yaml's search_tools block wires the LAN's SearXNG instance in as a standalone REST endpoint, not a model-callable tool — call it directly:

curl http://<proxy>:4000/v1/search/searxng-search \
  -H "Authorization: Bearer <a virtual key>" \
  -H "Content-Type: application/json" \
  -d '{"query": "...", "max_results": 5}'

Because this doesn't ask the model to emit a tool call, it sidesteps Qwen3.8-27B's known-flaky tool-calling (docs/research/qwen3.8-27b-tool-calling.md) entirely. Open WebUI's own web-search setting can point at this endpoint the same way.

Requires SEARXNG_LAN_IP set in .env (SearXNG's stable LAN IP — use a static DHCP reservation) so the litellm container can resolve search.home via extra_hosts. Full research: docs/research/litellm-searxng-search.md.

Knowledgebase (vector store / RAG)

LiteLLM's native knowledgebase feature has no Qdrant backend — the qdrant service in this stack only serves Open WebUI's own separate RAG/Memory feature and is unrelated to this. The only self-hosted path is litellm-pgvector, a companion service backed by its own Postgres+pgvector database (pgvector-db), which this stack now runs alongside litellm. Full research: docs/research/litellm-knowledgebase.md.

New pieces:

  • embedding-server — a second llama.cpp instance (small footprint, nomic-embed-text-v1.5) serving /v1/embeddings. The chat model isn't embedding-trained and llama.cpp serves one model per process, so this can't just be a flag on llama-server.
  • pgvector-db — Postgres with the pgvector extension, separate from litellm-db.
  • litellm-pgvector — the connector service; built straight from its upstream repo (no published image exists).
  • litellm-config.yaml's local-embedding model entry and vector_store_registry block, tying it together.

First-time setup

docker compose --profile tools run --rm downloader-embedding   # fetch the embedding model
docker compose up -d embedding-server pgvector-db litellm-pgvector

Create a litellm-pgvector virtual key in LiteLLM's Admin UI (per docs/proxy-key-onboarding.md) and set it as LITELLM_PGVECTOR_EMBEDDING_KEY in .env — the connector calls back into litellm for embeddings, same as any other workload.

Loading memory into it

data/memory.md and data/claude-legacy-memory.md — Claude-memory-style fact files — get loaded via:

./scripts/ingest-memory.sh

One chunk per fact/paragraph line, tagged with source/section metadata. Re-run after editing either file (see the script's header comment for the no-dedup caveat).

Querying it

Via the OpenAI Assistants-style file_search tool on a chat completion:

{
  "model": "qwen3.8-27b-local",
  "messages": [...],
  "tools": [{"type": "file_search", "vector_store_ids": ["memory-and-notes"]}]
}

or directly: POST /v1/vector_stores/memory-and-notes/search with {"query": "..."}.