mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-11 12:11:45 +00:00
docs(sandbox): correct the parsing worker's GPU note
The sentence dates from when DOCLING_OCR_ENABLED=true meant docling plus RapidOCR. Under the new defaults OCR_ENABLED=true resolves to the native backend with OCR_ENGINE=tesseract, a CPU subprocess that GPU libraries do not accelerate. Name what actually uses a GPU: docling's torch-backed layout and table models on the worker, or a DeepSeek-OCR endpoint off it.
This commit is contained in:
1 parent
b473498baf
commit
5a80e8cf22
1 file changed
+7
-2
@@ -202,8 +202,13 @@ Run a dedicated parsing worker that consumes the `parsing` queue:
|
||||
celery -A application.app.celery worker -Q parsing -l INFO
|
||||
```
|
||||
|
||||
It can be GPU-enabled with its own env (`OCR_ENABLED=true` plus GPU
|
||||
libraries) so OCR-heavy parsing runs on a separate, optionally larger pool.
|
||||
It takes its own env, so parse-heavy work runs on a separate, optionally larger
|
||||
pool. A GPU helps it only with the docling extra installed
|
||||
(`INSTALL_DOCLING=true`), whose layout and table models run on torch:
|
||||
`OCR_ENABLED=true` alone keeps the CPU-only native backend with the default
|
||||
`OCR_ENGINE=tesseract`, which GPU libraries do not accelerate. Setting
|
||||
`OCR_ENGINE=deepseek` instead moves the OCR cost onto the Ollama/vLLM endpoint
|
||||
and leaves this worker light.
|
||||
|
||||
**Dev / single-worker setups:** without a dedicated parsing worker the default
|
||||
worker must also consume `parsing`, or the tool's await never resolves:
|
||||
|
||||
Reference in new issue
Block a user