Backend (arc53/docsgpt): 4.5 GB compressed -> 0.9 GB with both embedding
models and tiktoken baked in.
- torch/transformers gone from the default install (docling extra only).
- Ubuntu 24.04 ships python3.12: no deadsnakes PPA, no software-properties-
common; every pin is a wheel, so no gcc/g++/rust in the builder.
- COPY --chown and a prefetch that runs as the process user replace the
trailing chown -R, which duplicated the 600 MB model layer.
- .dockerignore keeps __pycache__, .coverage, local indexes and .env out.
- EXTRAS build arg (INSTALL_DOCLING kept as an alias); the docling variant
also bakes docling's layout/table/RapidOCR models (DOCLING_ARTIFACTS_PATH)
and tesseract, and drops only the discovery documents of Google APIs the
app never builds.
- FLASK_DEBUG env removed (unused); OCI labels added.
Frontend (arc53/docsgpt-fe): 302 MB Vite dev server -> 25 MB static build
behind nginx. VITE_* variables are injected at container start into
/config.js and read through src/env.ts, so the image no longer needs a
rebuild per deployment; docker-compose.yaml keeps hot reload via the dev
target.
Publishing: every release and develop build now pushes a slim tag and a
-docling tag (docling engine + models + tesseract). docker-compose-hub.yaml
takes DOCSGPT_IMAGE_TAG / DOCSGPT_IMAGE_VARIANT; docker-compose-standalone.yaml
runs the stack from pre-built images without a checkout and is attached to
each release. setup.sh selects the -docling variant for OCR instead of
requiring a local build. A new workflow builds the image on PRs that touch
it and runs verify_offline under --network none; lint checks the exported
requirements match uv.lock.
Conflicts, and how each was taken:
- application/core/settings.py — ours. The renamed OCR_ENABLED /
OCR_ATTACHMENTS_ENABLED / OCR_MIN_CHARS_PER_PAGE accept main's
DOCLING_OCR_* spellings as AliasChoices, so nothing is dropped.
- application/Dockerfile — both. Main's install layers plus the
INSTALL_DOCLING build arg.
- application/parser/file/constants.py — both imports.
- deployment/docker-compose.yaml — both. The INSTALL_DOCLING /
INSTALL_TESSERACT build args on backend and worker, and main's
-Q docsgpt,parsing,embeddings, which query embedding needs.
- tests/conftest.py — theirs. Both sides fixed the same pytest-postgresql
9.0.0 autocommit= breakage; main's spelling is the one already on main.
- application/requirements.txt — the comments claimed different reasons
torch is in core. Main's is the true one now: it removed
sentence-transformers, so docling is torch's only remaining consumer.
Two things the merge broke without conflicting:
- onnxruntime. This branch moved it out of core into the docling extra;
main meanwhile made it the runtime local embeddings execute on
(fastembed). Git took the deletion, leaving fastembed with no pinned
runtime in a repo that pins everything. Restored to core, and no longer
pinned twice from the extra.
- The frontend copy of ATTACHMENT_PARSER_EXTENSIONS. The backend list is
derived and picked up the anydoc suffixes; the hand-kept frontend mirror
did not, so the composer would refuse files the API accepts.
tests/parser/file/test_constants.py is what caught it.
ruff, pytest (9897 passed), frontend build and docs build all pass. The
image build is unverified: no Docker daemon on this machine.
Both setup scripts start with `compose pull && compose up -d`. In the local
compose file backend and worker are build-only services, so `up -d` builds
only when no image exists yet: a rerun that switches OCR on wrote
INSTALL_TESSERACT=true to .env and then reused the image built without it,
leaving OCR_ENABLED=true with no engine. Build explicitly on the local
compose path; the hub path stays pull-only, its services have no build stage.
The OCR message named OCR_ENGINE=deepseek but not OCR_DEEPSEEK_URL, whose
default (localhost:11434) resolves to the container, not the host.
deployment/sandbox/README.md still described Docling as already present in
application/requirements.txt.
- Replay prior-turn assistant text as input_text, not output_text: a
Responses easy-input message only accepts input_* content parts, so the
old shape would 400 on the second turn of every conversation (default
store=false resends prior turns inline).
- Always request include=["reasoning.encrypted_content"] so in-turn
reasoning carryover works whether or not the response is also stored
server-side (OPENAI_RESPONSES_STORE).
- Add detail="auto" to input_image parts.
- Remove the now-dead azure_openai option from setup.sh / setup.ps1 and
the docs, completing the AzureOpenAILLM removal.
- Tests: correct the input_text / image-detail / include assertions and
add parallel tool-call and stream-error coverage.
* fixes setup scripts
fixes to env handling in setup script plus other minor fixes
* Remove var declarations
Declarations such as `LLM_PROVIDER=$LLM_PROVIDER` override .env variables in compose
Similar issue is present in the frontend - need to choose either to switch to separate frontend env or keep as is.
* Manage apikeys in settings
1. More pydantic management of api keys.
2. Clean up of variable declarations from docker compose files, used to block .env imports. Now should be managed ether by settings.py defaults or .env
This script includes the necessary changes to use container linking and updated environment variables for the `backend` and `worker` containers.
Make sure you have the `./frontend` and `./application` directories in the correct locations before running the script.