diff --git a/docs/research/langchain-pgvector-vs-litellm-pgvector.md b/docs/research/langchain-pgvector-vs-litellm-pgvector.md new file mode 100644 index 0000000..e2e2d48 --- /dev/null +++ b/docs/research/langchain-pgvector-vs-litellm-pgvector.md @@ -0,0 +1,388 @@ +# Research: LangChain+pgvector-direct RAG vs. the litellm-pgvector connector (for Gitea issue #25) + +**Question:** Is a LangChain + pgvector-direct retrieval architecture (LangChain's +`PGVector` class talking to Postgres directly, LiteLLM used only for +provider-abstracted chat/embedding calls) a better fit for this stack than the +already-built `litellm-pgvector` connector service (`vendor/litellm-pgvector/`, +issue #23's resolution)? + +**Bottom line: yes — switch.** Replace `litellm-pgvector` / +`vendor/litellm-pgvector/` with a small hand-written FastAPI service (roughly +80–120 lines, one new dependency set: `langchain-postgres`, `langchain-openai`, +`fastapi`, `uvicorn`) that wraps LangChain's `PGVector` class and calls +LiteLLM's OpenAI-compatible `/v1/embeddings` endpoint for embeddings. This +keeps "memory served from the gateway layer" in spirit (LiteLLM still does all +chat/embedding model calls; only the retrieval *query* now goes through a +LangChain-based service instead of a vendored 754-line Prisma-backed +connector), removes 754 lines of unvetted vendored source and a Prisma +migration dependency, and — critically — removes the *guessed and unverified* +`vector_store_registry` / `custom_llm_provider: pg_vector` config that issue +#24 flagged as never confirmed against a live deploy. `pgvector-db` (plain +`pgvector/pgvector:pg16`) is reusable completely as-is; only the schema +management strategy changes (LangChain manages its own two tables instead of +Prisma migrations). + +## 1. Verifying the markaicode article against LangChain's actual docs + +The article () proposes +three services — LangServe (port 8000, wraps a LangChain chain as REST), +LiteLLM (port 4000), and Postgres+pgvector (port 5432) — and shows: + +```python +vectorstore = PGVector( + connection_string="postgresql://user:pass@pgvector:5432/vectordb", + embedding_function=your_embeddings, + collection_name="documents", +) +``` + +**This snippet uses the deprecated class.** `connection_string` and +`embedding_function` are constructor args of +`langchain_community.vectorstores.pgvector.PGVector`, the older integration +that ships inside the general-purpose `langchain-community` package. The +current, actively-maintained integration lives in its own package, +`langchain-postgres`, as `langchain_postgres.vectorstores.PGVector`, with a +different constructor: + +```python +from langchain_postgres import PGVector +from langchain_openai import OpenAIEmbeddings + +vector_store = PGVector( + embeddings=embeddings, # not embedding_function + collection_name="my_docs", + connection="postgresql+psycopg://user:pass@host:5432/db", # not connection_string; psycopg3 DSN + use_jsonb=True, +) +``` + +Confirmed constructor signature (`embeddings`, `connection`, `collection_name`, +`use_jsonb`, `create_extension`, `async_mode`) via +[reference.langchain.com/python/langchain-postgres/vectorstores/PGVector](https://reference.langchain.com/python/langchain-postgres/vectorstores/PGVector) +and the setup walkthrough at +[docs.langchain.com/oss/python/integrations/vectorstores/pgvector](https://docs.langchain.com/oss/python/integrations/vectorstores/pgvector). +Two concrete corrections to the article's shape: + +- **Package**: `langchain-postgres` is a separate pip install + (`pip install -qU langchain-postgres`), not part of `langchain_community`. + The docs page does not say the community class is formally deprecated, but + it is the older/legacy path — the actively-documented integration page for + Postgres+pgvector points at `langchain_postgres.vectorstores.PGVector`. +- **Driver**: `langchain-postgres` requires the `psycopg` (psycopg3) driver + and a `postgresql+psycopg://` DSN, not `psycopg2` / plain `postgresql://`. + This repo's `pgvector-db` service doesn't care (it's just Postgres), but the + new service's `requirements.txt` needs `psycopg[binary]`, not + `psycopg2-binary` (which `vendor/litellm-pgvector/requirements.txt` uses). +- **`create_extension`** defaults to `True` — `PGVector` runs `CREATE + EXTENSION IF NOT EXISTS vector` itself on first connect. Source: same API + reference page above. This is the same requirement `litellm-pgvector`'s + Prisma schema has (Postgres user needs `CREATE` privilege) — **no change in + privilege requirements**, just who issues the DDL. +- The article's LangServe wrapper (`add_routes()`) is a real LangChain + component but adds a fourth abstraction (LangServe app → LangChain + chain/retriever → PGVector → Postgres) for what this stack needs from only + two HTTP endpoints (ingest, query). A plain FastAPI app calling `PGVector` + methods directly is simpler and is what's specced below (§5) — LangServe is + overkill for a two-endpoint internal service. + +There is also a v2, engine-based API, `langchain_postgres.v2.vectorstores.PGVectorStore` +(with native async via `PGEngine`/`AsyncPGVectorStore`), documented at +[reference.langchain.com/python/integrations/langchain_postgres](https://reference.langchain.com/python/integrations/langchain_postgres/). +It's newer and adds first-class async, but the plain `PGVector` class is the +one shown on LangChain's main pgvector integration page, is simpler to set up +(no separate `PGEngine` object), and this service has no need for the async +throughput `PGVectorStore` targets (single ingest script, occasional queries). +**Recommendation: use `langchain_postgres.vectorstores.PGVector`, not +`PGVectorStore`** — revisit only if query volume ever demands async +throughput. + +## 2. Complexity comparison: vendor/litellm-pgvector vs. a LangChain-direct service + +`vendor/litellm-pgvector/` as it exists in this repo today: + +| File | Lines | +|---|---| +| `main.py` | 521 | +| `embedding_service.py` | 89 | +| `models.py` | 85 | +| `config.py` | 59 | +| `prisma/schema.prisma` | 39 | +| **Total (Python + schema)** | **793** | + +Plus: a `Dockerfile` that runs `prisma generate` at build time, a +`requirements.txt` pulling in `prisma==0.11.0` (which itself vendors a Rust +query-engine binary download at install time — a second network dependency +inside the Docker build, on top of the git-context build that already failed +once per the vendoring rationale in `docker-compose.yml`/`VENDORED.md`), and +hand-rolled raw-SQL string building for every endpoint (see `main.py` lines +86–128, 219–320, 322–402, 405–511 — table/column names are +f-string-interpolated from `settings.table_names`/`settings.db_fields`, not +parameterized, though the *values* are parameterized). None of this has been +smoke-tested against a live deploy (issue #24). + +A LangChain-direct replacement needs to expose exactly two HTTP endpoints +(`POST /ingest`, `POST /query` — see §5) as a thin wrapper around library +calls that already do the embedding + SQL work: + +- `langchain_postgres.PGVector.add_texts(texts, metadatas)` — ingestion, + replaces `main.py`'s `create_embedding`/`create_embeddings_batch` (176 + lines) and all of `models.py`'s embedding request/response schemas. +- `langchain_postgres.PGVector.similarity_search_with_relevance_scores(query, + k)` — query, replaces `search_vector_store` (100 lines) and + `generate_query_embedding`/`embedding_service.py` (89 lines) — LangChain's + `OpenAIEmbeddings` (from `langchain-openai`, pointed at + `base_url=http://litellm:4000/v1`) does the embedding call instead of a + hand-written `httpx` call to LiteLLM. +- No `create_vector_store`/`list_vector_stores` needed — one Postgres + "collection" (LangChain's term, one row in its `langchain_pg_collection` + table) is enough for this repo's single `memory-and-notes` use case; + `collection_name="memory-and-notes"` passed once at construction covers it. +- No Prisma, no generated query-engine binary, no `prisma/schema.prisma` — + `PGVector` creates and manages its own two tables + (`langchain_pg_collection`, `langchain_pg_embedding`) via SQLAlchemy on + first use. + +Estimated replacement size: **~90–120 lines of Python** across the FastAPI app ++ Pydantic request/response models (a `/ingest` and `/query` endpoint, an +`Embeddings`-backed `PGVector` instance built once at startup, an API-key +`Depends()` check reused almost verbatim from `vendor/litellm-pgvector/main.py` +lines 50–55) — roughly an **85% reduction** from the vendored connector's 793 +lines, with zero Prisma/migration surface and no hand-built raw SQL (parameter +binding, vector-string formatting, and JSONB metadata handling are all +internal to `PGVector`, not re-implemented per endpoint as in `main.py`). + +**This does not eliminate the need for a small custom service** — a bare +`PGVector.similarity_search()` call is a Python library call, not an HTTP +endpoint, and this stack's callers (the ingest script, Claude Code hooks +described in `docs/memory-knowledgebase.md`) need HTTP. Something still has to +be that thin wrapper; it's just ~85% smaller and has no ORM/migration layer. + +## 3. Does this still count as "memory served from the gateway layer"? + +Yes, with the same division of responsibility issue #21 established, just a +narrower one for the connector. LiteLLM (`litellm` container) remains the +**only** thing every client (Open WebUI, Claude Code, anything else behind the +proxy) talks to for chat and embedding *model* calls — nothing changes there. +What moves is which service holds the *retrieval* logic: today it's +`litellm-pgvector` (itself calling back into `litellm:4000` for embeddings, +per `docker-compose.yml` lines 241–246); under this alternative it's the new +LangChain-based service, which **also** calls back into `litellm:4000/v1` for +embeddings (via `langchain-openai`'s `OpenAIEmbeddings(base_url="http://litellm:4000/v1")`, +since LiteLLM exposes an OpenAI-compatible API — confirmed by this repo's +existing `local-embedding` `model_list` entry in `litellm-config.yaml` already +being OpenAI-route-compatible). The client-facing shape is identical: no +client (Open WebUI, a Claude Code hook, `ingest-memory.sh`) talks to Postgres +directly, and no client bypasses LiteLLM for the actual embedding/generation +model calls. Only the internal retrieval connector's implementation changes, +from a vendored third-party FastAPI+Prisma app to an in-repo FastAPI+LangChain +app. This does not contradict #21's standing decision. + +One nuance: today, `vector_store_registry` makes LiteLLM itself the thing a +chat completion's `tools: [{"type": "file_search", ...}]` call reaches (per +`docs/research/litellm-knowledgebase.md` §1) — LiteLLM intercepts the tool +call and proxies to the registered `pg_vector` backend. **Under the +alternative, that specific mechanism goes away**: the new service is not a +`vector_store_registry` provider (there is no `langchain_postgres`/custom +provider type in LiteLLM's `vector_store_registry`, and building one would +reintroduce exactly the custom-connector complexity this ticket is trying to +shed). Retrieval becomes a direct call to the new service's `/query` endpoint +(or a client-side RAG step: query the service, splice results into the prompt +before calling `litellm`) rather than an in-band `file_search` tool call +LiteLLM resolves itself. This is a real, if modest, capability reduction: +`{"type": "file_search", "vector_store_ids": [...]}` inside a `/chat/completions` +call to `litellm` stops working; callers move to a two-step "search then +send" pattern instead. Given `vector_store_registry`'s exact field shape was +never confirmed working (issue #24) and this is the only capability lost, +this is judged an acceptable trade for the complexity and reliability win. + +## 4. Ingestion, embedding-server, pgvector-db reuse + +**`embedding-server`** (llama.cpp, `nomic-embed-text-v1.5`, `docker-compose.yml` +lines 38–66): fully reusable as-is. It already serves an OpenAI-compatible +`/v1/embeddings` route (llama.cpp's `--embeddings` flag, per +`docs/research/litellm-knowledgebase.md` §3), and `litellm-config.yaml` +already exposes it through LiteLLM as the `local-embedding` model +(`api_base: http://embedding-server:8080/v1`). The new service reaches it the +same way `litellm-pgvector` does today — indirectly, by calling +`http://litellm:4000/v1/embeddings` with `model=local-embedding` — via +`langchain_openai.OpenAIEmbeddings(model="local-embedding", +base_url="http://litellm:4000/v1", api_key=)`. `OpenAIEmbeddings` +is LangChain's standard `Embeddings` implementation for any OpenAI-compatible +endpoint (no custom `Embeddings` subclass needed) — confirmed via the +constructor example in +[docs.langchain.com/oss/python/integrations/vectorstores/pgvector](https://docs.langchain.com/oss/python/integrations/vectorstores/pgvector), +which pairs `OpenAIEmbeddings` with `PGVector` directly. + +**`scripts/ingest-memory.sh`** changes, but modestly: same per-line chunking +of `data/memory.md`/`data/claude-legacy-memory.md` (the ponytail comment in +the current script — one chunk per non-empty, non-heading line, no real +chunking logic needed — still holds), but it now `curl`s the new service's +`POST /ingest` instead of `litellm-pgvector`'s two-step +`POST /v1/vector_stores` + `POST /v1/vector_stores/{id}/embeddings/batch`. +Concretely: build one JSON array of `{content, metadata}` per file (same +shape the script already builds) and POST it once to `/ingest`; the new +service calls `PGVector.add_texts(texts, metadatas)` internally, which +embeds and inserts in one call — no separate "create the store" step, since +`collection_name` is fixed at service startup and `PGVector` creates the +collection row implicitly on first `add_texts`. The no-dedup caveat in the +current script's header comment carries over unchanged (`add_texts` always +inserts; nothing in `PGVector` deduplicates by content). + +**`pgvector-db`** (`pgvector/pgvector:pg16`, `docker-compose.yml` lines +198–214): reusable completely as-is — same image, same "separate Postgres +instance from `litellm-db`" rationale, no compose changes needed to the +service definition itself. Only the *schema* changes: `litellm-pgvector`'s +Prisma schema (`vendor/litellm-pgvector/prisma/schema.prisma`, 39 lines) +defines its own `vector_stores`/embeddings tables with configurable +table/field names (`config.py`'s `settings.table_names`/`db_fields`), applied +via Prisma migrations at container startup (the exact behavior of which, +per issue #24's ponytail note, was never smoke-tested). `PGVector` instead +auto-creates two fixed tables — `langchain_pg_collection` and +`langchain_pg_embedding` — via SQLAlchemy `Base.metadata.create_all()`-style +setup on first connection, no separate migration step or migration tool +needed. This is a strictly simpler schema-management story: no Prisma CLI, +no `prisma generate` build step, no migration files to track. The existing +`litellm_pgvector`/`PGVECTOR_DB_PASSWORD` database and role stay — the new +service just needs its own `DATABASE_URL` pointed at the same database (or a +fresh one; reusing `litellm_pgvector` is simplest since no data has been +proven-loaded into it yet per issue #24). + +**`litellm-pgvector` (container) and `vendor/litellm-pgvector/` (source +tree)**: removed if this verdict is adopted. Nothing else in the repo depends +on the vendored source once the new service replaces it. + +## 5. Concrete replacement spec (not implemented — spec only) + +**New compose service**, replacing the `litellm-pgvector` block: + +```yaml + memory-retrieval: # name TBD; "litellm-pgvector" freed up + build: + context: ./services/memory-retrieval # new small in-repo dir, not vendored + container_name: memory-retrieval + depends_on: + pgvector-db: + condition: service_healthy + litellm: + condition: service_healthy + environment: + - DATABASE_URL=postgresql+psycopg://litellm_pgvector:${PGVECTOR_DB_PASSWORD}@pgvector-db:5432/litellm_pgvector + - LITELLM_BASE_URL=http://litellm:4000/v1 + - LITELLM_API_KEY=${MEMORY_RETRIEVAL_EMBEDDING_KEY} # new virtual key, same pattern as LITELLM_PGVECTOR_EMBEDDING_KEY today + - EMBEDDING_MODEL=local-embedding + - SERVER_API_KEY=${MEMORY_RETRIEVAL_API_KEY} # replaces LITELLM_PGVECTOR_API_KEY + - COLLECTION_NAME=memory-and-notes + expose: + - "8000" + restart: unless-stopped + networks: [ai-stack] +``` + +Note `DATABASE_URL` changes scheme to `postgresql+psycopg://` (psycopg3 DSN +`langchain-postgres` requires — see §1), not the `postgresql://` scheme +`litellm-pgvector` uses today. + +**Dockerfile shape** (`services/memory-retrieval/Dockerfile`): same base image +family as today's connector — `python:3.11-slim` — `pip install` the +requirements, copy the app, `CMD ["uvicorn", "main:app", "--host", "0.0.0.0", +"--port", "8000"]`. No `prisma generate` step, no `build-essential`/`postgresql-client` +needed at the OS-package level (no Prisma binary to build against; `psycopg[binary]` +ships prebuilt wheels). + +**`requirements.txt`**: `fastapi`, `uvicorn[standard]`, `langchain-postgres`, +`langchain-openai`, `psycopg[binary]`, `pydantic` — five packages vs. the +current ten in `vendor/litellm-pgvector/requirements.txt`, and no `prisma`. + +**Endpoints** (2, vs. today's 5 — create store, list stores, search, create +embedding, create embeddings batch): + +- `POST /ingest` — body: `{"chunks": [{"content": "...", "metadata": {...}}]}`. + Calls `PGVector.add_texts(texts=[...], metadatas=[...])`. Replaces both + `ingest-memory.sh`'s "create store" call (no longer needed — collection is + implicit) and its batch-embeddings call. +- `POST /query` — body: `{"query": "...", "k": 5}`. Calls + `PGVector.similarity_search_with_relevance_scores(query, k=k)`, returns + `{"results": [{"content": ..., "metadata": ..., "score": ...}]}`. +- `GET /health` — same trivial liveness check the current connector has + (`main.py` lines 514–517), kept for compose `depends_on`/`healthcheck` use + if desired. + +Bearer-token auth via `SERVER_API_KEY`, same `Depends()` pattern as +`vendor/litellm-pgvector/main.py` lines 47–55 — that part of the existing +connector is fine as-is and can be copied over close to verbatim. + +**`litellm-config.yaml` change**: remove the `vector_store_registry` block +entirely (lines 47–61 today) — per §3, the new service isn't a +`vector_store_registry` provider, so there's nothing to register. The +`local-embedding` `model_list` entry (lines 25–35) is unchanged and still +needed — it's what the new service's `OpenAIEmbeddings` client calls through +`litellm`. + +## 6. Verdict + +**Switch to LangChain-direct.** This does invalidate part of issue #23's +resolution and the `litellm-pgvector` container currently in +`docker-compose.yml` — say so explicitly for scoping a follow-up +implementation ticket: + +- **Torn out**: `litellm-pgvector` service block in `docker-compose.yml`, + `vendor/litellm-pgvector/` (entire tree), the `vector_store_registry` block + in `litellm-config.yaml`, `LITELLM_PGVECTOR_API_KEY`/ + `LITELLM_PGVECTOR_EMBEDDING_KEY` env vars (replaced by new-service + equivalents). +- **Survives unchanged**: `pgvector-db` (compose block as-is), `embedding-server` + (compose block as-is), the `local-embedding` `model_list` entry in + `litellm-config.yaml`, the overall "LiteLLM does all model calls" gateway + boundary from issue #21. +- **New**: a small in-repo `services/memory-retrieval/` FastAPI+LangChain + app (~90–120 lines, spec in §5), a new compose service block, a rewritten + `scripts/ingest-memory.sh` pointed at `/ingest` instead of + `litellm-pgvector`'s two-step API, and one new virtual key + (`MEMORY_RETRIEVAL_EMBEDDING_KEY`) issued the same way + `LITELLM_PGVECTOR_EMBEDDING_KEY` was per `docs/proxy-key-onboarding.md`. +- **Capability trade-off** (§3): the OpenAI Assistants-style `file_search` + tool call directly against `litellm`'s `/chat/completions` stops working + (no `vector_store_registry` entry to resolve it); callers move to querying + the new service directly and splicing results into the prompt themselves. + This mechanism was never confirmed working in the first place (issue #24), + so it's a loss of unverified capability, not working functionality. + +This is not implemented here — this document is research/spec only, per the +ticket's scope. A follow-up implementation ticket should: write +`services/memory-retrieval/`, update `docker-compose.yml` and +`litellm-config.yaml` per §5, rewrite `scripts/ingest-memory.sh`, delete +`vendor/litellm-pgvector/`, and update `docs/memory-knowledgebase.md`. + +## 7. Repo context read for this research + +- `docker-compose.yml` — `embedding-server`, `pgvector-db`, `litellm-pgvector` + service definitions (build context, env vars, network). +- `litellm-config.yaml` — `local-embedding` model entry, `vector_store_registry` + block. +- `docs/research/litellm-knowledgebase.md` — prior research establishing + LiteLLM's native vector-store feature has no Qdrant provider and that + `litellm-pgvector` was the only self-hosted path found at the time; its §1, + §3, §4 are cited above for the `file_search` tool-call mechanism and + embedding-model wiring. +- `docs/memory-knowledgebase.md` — current user-facing setup doc for the + knowledgebase/memory feature. +- `scripts/ingest-memory.sh` — current ingestion mechanism and its no-dedup + caveat. +- `vendor/litellm-pgvector/{main.py,embedding_service.py,models.py,config.py, + prisma/schema.prisma,requirements.txt,Dockerfile}` — read in full to + characterize size (793 total lines Python+schema) and shape (raw-SQL + f-string-built queries per endpoint, Prisma-managed schema, calls back to + LiteLLM's `/embeddings` for query embedding). +- Gitea issue #25 (this ticket) and its parent #21 — standing decision that + memory is served from the gateway layer, not a client. +- Gitea issue #23/#24 (referenced by #25) — original `litellm-pgvector` + adoption and its unverified-against-live-deploy caveat. + +## 8. Primary sources consulted + +- [markaicode.com/stack/litellm-langchain-stack/](https://markaicode.com/stack/litellm-langchain-stack/) — the alternative surfaced by the user; verified and corrected above (§1). +- [reference.langchain.com/python/langchain-postgres/vectorstores/PGVector](https://reference.langchain.com/python/langchain-postgres/vectorstores/PGVector) — `PGVector` constructor signature, `create_extension` default. +- [docs.langchain.com/oss/python/integrations/vectorstores/pgvector](https://docs.langchain.com/oss/python/integrations/vectorstores/pgvector) — install instructions, `postgresql+psycopg://` DSN requirement, `OpenAIEmbeddings` + `PGVector` pairing example. +- [github.com/langchain-ai/langchain-postgres](https://github.com/langchain-ai/langchain-postgres) — package README; async-function-pairing note; `PGVector` deprecated in v0.0.14+ in favor of `PGVectorStore` (v2 API). +- [reference.langchain.com/python/integrations/langchain_postgres/](https://reference.langchain.com/python/integrations/langchain_postgres/) — v2 `PGVectorStore`/`PGEngine`/`AsyncPGVectorStore` API, compared against `PGVector` in §1. +- `vendor/litellm-pgvector/README.md` and source files (this repo) — connector's existing endpoint/config shape, cited for the LOC comparison in §2.