docs(sandbox): correct the parsing worker's GPU note

The sentence dates from when DOCLING_OCR_ENABLED=true meant docling plus
RapidOCR. Under the new defaults OCR_ENABLED=true resolves to the native
backend with OCR_ENGINE=tesseract, a CPU subprocess that GPU libraries do not
accelerate. Name what actually uses a GPU: docling's torch-backed layout and
table models on the worker, or a DeepSeek-OCR endpoint off it.
This commit is contained in:
Alex committed 2026-09-04 16:02:50 +01:00
1 parent b473498baf
commit 5a80e8cf22
1 file changed
+7 -2
+7 -2
View File
@@ -202,8 +202,13 @@ Run a dedicated parsing worker that consumes the `parsing` queue:
celery -A application.app.celery worker -Q parsing -l INFO
```
It can be GPU-enabled with its own env (`OCR_ENABLED=true` plus GPU
libraries) so OCR-heavy parsing runs on a separate, optionally larger pool.
It takes its own env, so parse-heavy work runs on a separate, optionally larger
pool. A GPU helps it only with the docling extra installed
(`INSTALL_DOCLING=true`), whose layout and table models run on torch:
`OCR_ENABLED=true` alone keeps the CPU-only native backend with the default
`OCR_ENGINE=tesseract`, which GPU libraries do not accelerate. Setting
`OCR_ENGINE=deepseek` instead moves the OCR cost onto the Ollama/vLLM endpoint
and leaves this worker light.
**Dev / single-worker setups:** without a dedicated parsing worker the default
worker must also consume `parsing`, or the tool's await never resolves: