Today's embedding-server crash-loop (missing nomic-embed-text GGUF) was a manual step nobody ran. update.sh now runs both downloader profiles itself, every time, before bringing services up -- no separate command to remember. - docker-compose.yml: downloader/downloader-embedding commands gain a `test -f ... && skip || curl ...` guard, so re-running update.sh never re-downloads an existing model file. - scripts/update.sh: runs both profiles after image pull/build, before service recreation. - scripts/download-model.sh removed -- folded in, redundant standalone script. - README.md / docs/memory-knowledgebase.md updated accordingly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WHfjWrSEcGhCoeu6dQfDa
67 lines
4.0 KiB
Markdown
67 lines
4.0 KiB
Markdown
# Knowledgebase, memory, and web search
|
|
|
|
Three gateway-level capabilities added on top of the [AI gateway/proxy](https://git.arthurerlich.de/haylan/LLM-Server/issues/9), so every client behind LiteLLM gets them — not just Open WebUI. See [issue #21](https://git.arthurerlich.de/haylan/LLM-Server/issues/21) for the rationale.
|
|
|
|
**Not yet verified on real hardware** — see [issue #24](https://git.arthurerlich.de/haylan/LLM-Server/issues/24). In particular: `litellm-pgvector`'s Prisma migrations on first boot, and the exact `vector_store_registry` field names for the `pg_vector` provider.
|
|
|
|
## Web search (SearXNG)
|
|
|
|
`litellm-config.yaml`'s `search_tools` block wires the LAN's SearXNG instance in as a **standalone REST endpoint**, not a model-callable tool — call it directly:
|
|
|
|
```bash
|
|
curl http://<proxy>:4000/v1/search/searxng-search \
|
|
-H "Authorization: Bearer <a virtual key>" \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"query": "...", "max_results": 5}'
|
|
```
|
|
|
|
Because this doesn't ask the model to emit a tool call, it sidesteps Qwen3.8-27B's known-flaky tool-calling (`docs/research/qwen3.8-27b-tool-calling.md`) entirely. Open WebUI's own web-search setting can point at this endpoint the same way.
|
|
|
|
Requires `SEARXNG_LAN_IP` set in `.env` so the `litellm` container can resolve `search.home` via `extra_hosts` — `./scripts/update.sh` resolves and fills this in automatically from the host's own DNS if it's blank (use a static DHCP reservation for `search.home` so it doesn't drift). Full research: `docs/research/litellm-searxng-search.md`.
|
|
|
|
## Knowledgebase (vector store / RAG)
|
|
|
|
LiteLLM's native knowledgebase feature has **no Qdrant backend** — the `qdrant` service in this stack only serves Open WebUI's own separate RAG/Memory feature and is unrelated to this. The only self-hosted path is [litellm-pgvector](https://github.com/BerriAI/litellm-pgvector), a companion service backed by its own Postgres+pgvector database (`pgvector-db`), which this stack now runs alongside `litellm`. Full research: `docs/research/litellm-knowledgebase.md`.
|
|
|
|
New pieces:
|
|
|
|
- **`embedding-server`** — a second llama.cpp instance (small footprint, `nomic-embed-text-v1.5`) serving `/v1/embeddings`. The chat model isn't embedding-trained and llama.cpp serves one model per process, so this can't just be a flag on `llama-server`.
|
|
- **`pgvector-db`** — Postgres with the pgvector extension, separate from `litellm-db`.
|
|
- **`litellm-pgvector`** — the connector service; no published image exists, so it's built from a vendored copy of the upstream repo at `vendor/litellm-pgvector/` (see that dir's `VENDORED.md`) — a remote git build context failed on the server's Docker/BuildKit setup.
|
|
- `litellm-config.yaml`'s `local-embedding` model entry and `vector_store_registry` block, tying it together.
|
|
|
|
### First-time setup
|
|
|
|
`./scripts/update.sh` fetches the embedding model automatically (skips it if already downloaded). To do it by hand instead:
|
|
|
|
```bash
|
|
docker compose --profile tools run --rm downloader-embedding # fetch the embedding model
|
|
docker compose up -d embedding-server pgvector-db litellm-pgvector
|
|
```
|
|
|
|
`./scripts/update.sh` mints `LITELLM_PGVECTOR_EMBEDDING_KEY` automatically (a `litellm-pgvector` virtual key via LiteLLM's own API) if it's blank — it calls back into `litellm` for embeddings, same as any other workload. See `docs/proxy-key-onboarding.md` if a mint fails and it needs doing by hand.
|
|
|
|
### Loading memory into it
|
|
|
|
`data/memory.md` and `data/claude-legacy-memory.md` — Claude-memory-style fact files — get loaded via:
|
|
|
|
```bash
|
|
./scripts/ingest-memory.sh
|
|
```
|
|
|
|
One chunk per fact/paragraph line, tagged with `source`/`section` metadata. Re-run after editing either file (see the script's header comment for the no-dedup caveat).
|
|
|
|
### Querying it
|
|
|
|
Via the OpenAI Assistants-style `file_search` tool on a chat completion:
|
|
|
|
```json
|
|
{
|
|
"model": "qwen3.8-27b-local",
|
|
"messages": [...],
|
|
"tools": [{"type": "file_search", "vector_store_ids": ["memory-and-notes"]}]
|
|
}
|
|
```
|
|
|
|
or directly: `POST /v1/vector_stores/memory-and-notes/search` with `{"query": "..."}`.
|