mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-11 12:11:45 +00:00
feat: embed on the worker, and stop batching the ONNX pass
The API embeds every query it serves, so it held its own copy of the model: ~890 MB it never needed. EMBEDDINGS_DELEGATE_TO_WORKER (on by default) sends the text to the Celery worker instead and gets the vector back, taking an API process from 1176 MB to 285 MB with no ONNX Runtime imported at all. The client embeds locally when it finds itself inside a worker task, so the worker never dispatches to itself -- the same self-deadlock DOCUMENT_PARSE_QUEUE avoids on the parsing side. EMBEDDINGS_BASE_URL still wins over it, and remains the right answer for production. ensure_vector_schema was constructing the embeddings instance purely to read .dimension off it, loading several hundred MB of ONNX into every API and worker process at import. For a model the registry describes that is a lookup; only an unregistered name now falls back to loading. EMBEDDINGS_BATCH_SIZE was sizing two unrelated things: chunks per store transaction (and per remote embed request) and documents per ONNX forward pass. Each pass pads every input up to its longest, and that waste grows with the square of chunk length, so at the 1250-token default a batch of 32 peaked at 6.6 GB and took 326s where a batch of 1 peaked at 2.9 GB and took 90s. The forward pass is now sized by EMBEDDINGS_MODEL_BATCH_SIZE, defaulting to 1; storage and remote batching are unchanged at 32. reembed embeds in-process: a batch job that walks the whole index should not round-trip every chunk through a broker, and loading the model there reports a real failure instead of timing out against an empty queue. Also drops the mpnet zip download from the docs and the devcontainer, which pointed at a SentenceTransformers export with no ONNX graph and had been inert since the FastEmbed swap; corrects the claim that any sentence-transformers model works; and settles the Configuring/Settings pages on what the registry and the repository metadata actually decide.
This commit is contained in:
1 parent
ef51ee97b3
commit
47e53a71c4
20 files changed
+679
-266
No files matched your search
@@ -24,36 +24,30 @@ In essence, embedding models are the bridge that allows DocsGPT to understand th
|
||||
|
||||
DocsGPT is designed to be flexible and supports a wide range of embedding models right out of the box:
|
||||
|
||||
* **Sentence Transformers:** DocsGPT supports all models available through the [Sentence Transformers library](https://www.sbert.net/). This library offers a vast selection of pre-trained embedding models, known for their quality and efficiency in various semantic tasks. This is the default (`EMBEDDINGS_NAME=huggingface_sentence-transformers/all-mpnet-base-v2`).
|
||||
* **Local models (FastEmbed / ONNX Runtime):** DocsGPT runs local embeddings through [FastEmbed](https://github.com/qdrant/fastembed). Any of FastEmbed's built-in models works, as does any Hugging Face repository that ships an ONNX export at `onnx/model.onnx` — which covers most popular sentence-transformers repos. A repository with PyTorch weights only will not load; serve those over `EMBEDDINGS_BASE_URL` instead. New installs default to `ibm-granite/granite-embedding-311m-multilingual-r2`; existing ones stay on `huggingface_sentence-transformers/all-mpnet-base-v2` until re-embedded.
|
||||
* **OpenAI Embeddings:** DocsGPT supports OpenAI embedding models (for example `text-embedding-ada-002`, `text-embedding-3-small`, `text-embedding-3-large`) via the OpenAI API.
|
||||
* **Azure OpenAI Embeddings:** Set `AZURE_EMBEDDINGS_DEPLOYMENT_NAME` alongside your Azure OpenAI configuration.
|
||||
* **Remote OpenAI-compatible Embeddings:** Any server that exposes an OpenAI-compatible `/v1/embeddings` endpoint (for example llama.cpp, vLLM, TEI, or a hosted provider) by setting `EMBEDDINGS_BASE_URL`. See [Remote Embeddings](#remote-openai-compatible-embeddings) below.
|
||||
|
||||
## Configuring Sentence Transformer Models
|
||||
## Configuring a Local Model
|
||||
|
||||
To utilize Sentence Transformer models within DocsGPT, you need to follow these steps:
|
||||
Set `EMBEDDINGS_NAME` in your `.env` to a registry name or a Hugging Face repository id:
|
||||
|
||||
1. **Download the Model:** Sentence Transformer models are typically hosted on Hugging Face Model Hub. You need to download your chosen model and place it in the `model/` folder in the root directory of your DocsGPT project.
|
||||
```
|
||||
EMBEDDINGS_NAME=ibm-granite/granite-embedding-311m-multilingual-r2
|
||||
```
|
||||
|
||||
For example, to use the `all-mpnet-base-v2` model, you would set `EMBEDDINGS_NAME` as described below, and ensure that the model files are available locally (DocsGPT will attempt to download it if it's not found, but local download is recommended for development and offline use).
|
||||
The model is downloaded on first use and cached; set `EMBEDDINGS_CACHE_DIR` to control where. There is no `model/` folder to populate by hand, and a filesystem path is not accepted as a model name.
|
||||
|
||||
2. **Set `EMBEDDINGS_NAME` in `.env` (or `settings.py`):** You need to configure the `EMBEDDINGS_NAME` setting in your `.env` file (or `settings.py`) to point to the desired Sentence Transformer model.
|
||||
DocsGPT knows the pooling, vector width and context window of the models in its registry (`all-mpnet-base-v2`, `granite-embedding-311m-multilingual-r2`, `granite-embedding-97m-multilingual-r2`). For any other repository it reads those from the repository's own `1_Pooling/config.json` and `modules.json`. If a repository declares neither, mean pooling with L2 normalization is assumed and a warning is logged — pin the real values with `EMBEDDINGS_POOLING` (`cls` or `mean`) and `EMBEDDINGS_NORMALIZE`.
|
||||
|
||||
* **Using a pre-downloaded model from `model/` folder:** You can specify a path to the downloaded model within the `model/` directory. For instance, if you downloaded `all-mpnet-base-v2` and it's in `model/all-mpnet-base-v2`, you could potentially use a relative path like (though direct path to the model name is usually sufficient):
|
||||
Models with a Dense projection layer (for example `sentence-transformers/LaBSE`) are refused at startup: FastEmbed cannot apply the projection, so the vectors would be the wrong width and in a different space.
|
||||
|
||||
```
|
||||
EMBEDDINGS_NAME=huggingface_sentence-transformers/all-mpnet-base-v2
|
||||
```
|
||||
or simply use the model identifier:
|
||||
```
|
||||
EMBEDDINGS_NAME=sentence-transformers/all-mpnet-base-v2
|
||||
```
|
||||
For an offline or air-gapped install, pre-fetch the model at build or setup time:
|
||||
|
||||
* **Using a model directly from Hugging Face Model Hub:** You can directly specify the model identifier from Hugging Face Model Hub:
|
||||
|
||||
```
|
||||
EMBEDDINGS_NAME=huggingface_sentence-transformers/all-mpnet-base-v2
|
||||
```
|
||||
```bash
|
||||
python -m application.scripts.prefetch_models
|
||||
```
|
||||
|
||||
## Using OpenAI Embeddings
|
||||
|
||||
@@ -95,6 +89,41 @@ You usually do not need to set it. When `EMBEDDINGS_NAME` names a model DocsGPT
|
||||
|
||||
Leaving `EMBEDDINGS_NAME` unset imposes no limit: the name is only forwarded as the `model` field in each request, so a default nobody chose is not taken as a description of your server. When no model tokenizer is available, counting falls back to tiktoken — pick a limit with headroom below the server's true limit to absorb the skew between the two tokenizers.
|
||||
|
||||
## Where the model runs
|
||||
|
||||
A local embedding model costs a few hundred megabytes of resident memory per process, and the API embeds every query it serves — so by default it would hold its own copy alongside the worker's.
|
||||
|
||||
`EMBEDDINGS_DELEGATE_TO_WORKER` (on by default) moves that work to the Celery worker: the API sends the text over the broker and gets the vector back, holding no model. Measured on a default install, the API process drops from ~657 MB to ~284 MB, and query embedding costs one broker round trip (~60 ms on a prefork worker).
|
||||
|
||||
Retrieval then depends on a worker consuming `EMBEDDINGS_QUEUE` (`embeddings` by default). A bare `celery worker` with no `-Q` consumes it along with everything else, so the standard deployment works unchanged — but its concurrency is shared with ingest, so a query can queue behind a long parse. Run a dedicated worker to isolate query latency:
|
||||
|
||||
```bash
|
||||
celery -A application.app.celery worker -Q embeddings
|
||||
```
|
||||
|
||||
Set `EMBEDDINGS_DELEGATE_TO_WORKER=false` if you run the API without a worker; it will load the model in-process instead.
|
||||
|
||||
For production, prefer `EMBEDDINGS_BASE_URL`. A real embedding service removes the model from *both* the API and the worker, and replaces the broker round trip with a network call.
|
||||
|
||||
## Batch sizes
|
||||
|
||||
Two separate knobs, easily confused:
|
||||
|
||||
- `EMBEDDINGS_BATCH_SIZE` (default 32) — chunks per store transaction, and per request to a remote embeddings API. Larger means fewer round trips and fewer transactions.
|
||||
- `EMBEDDINGS_MODEL_BATCH_SIZE` (default 1) — documents per forward pass of a *local* model.
|
||||
|
||||
For the local model, bigger batches are not faster. ONNX needs a rectangular tensor, so every input in a pass is padded up to the longest one in it, and that waste grows with the square of chunk length. Measured on a 30-document ingest at the 1250-token default chunk size:
|
||||
|
||||
| `EMBEDDINGS_MODEL_BATCH_SIZE` | embed time | peak RSS |
|
||||
| --- | --- | --- |
|
||||
| 32 | 154 s | 7.7 GB |
|
||||
| 8 | 76 s | 5.0 GB |
|
||||
| 4 | 76 s | 3.6 GB |
|
||||
| 2 | 74 s | 2.3 GB |
|
||||
| 1 | 53 s | 1.5 GB |
|
||||
|
||||
Raise it only if your chunks are short and uniform in length.
|
||||
|
||||
## Important: Embedding Dimensions Must Stay Consistent
|
||||
|
||||
Each embedding model produces vectors of a fixed dimension, and your vector store is created with that dimension. **Changing `EMBEDDINGS_NAME` to a model with a different dimension is not compatible with an existing index** — FAISS and LanceDB will raise a dimension-mismatch error, and pgvector/Qdrant tables are sized to the original dimension.
|
||||
@@ -119,6 +148,6 @@ With `GRAPHRAG_ENABLED`, the script also rewrites `graph_nodes.name_embedding` o
|
||||
|
||||
## Adding Support for Other Embedding Models
|
||||
|
||||
If you wish to use an embedding model that is not supported out-of-the-box, a good starting point for adding custom embedding model support is to examine the `base.py` file located in the `application/vectorstore` directory.
|
||||
To teach DocsGPT about a new model — so it carries a known pooling, width and context window rather than being inferred — add an `EmbeddingModel` entry to `MODELS` in `application/vectorstore/model_registry.py`. That registry is the single source of truth the local runner, the remote client, the schema bootstrap and the chunker all read.
|
||||
|
||||
Specifically, pay attention to the `EmbeddingsWrapper` and `EmbeddingsSingleton` classes. `EmbeddingsWrapper` provides a way to wrap different embedding model libraries into a consistent interface for DocsGPT. `EmbeddingsSingleton` manages the instantiation and retrieval of embedding model instances. By understanding these classes and the existing embedding model implementations, you can create your own custom integration for virtually any embedding model library you desire.
|
||||
Reference in new issue
Block a user