Major bumps whose ceilings had to move. redis 8.1.0, tiktoken 0.14.0 and
daytona 0.211.2 needed no code change (the Daytona client, filesystem and
process signatures the sandbox calls are unchanged; redis 8 was checked
against a live server through the app's own sync and async clients).
reportlab 5.0.1 is test-only.
openapi-parser 2.0.0 is a rewrite onto pydantic spec models: `paths` is now
a dict keyed by URL rather than a list of objects carrying their own `url`,
and a path item exposes one field per HTTP method instead of an `operations`
list. `OpenAPI3Parser` reads both accordingly, iterating methods in the
order the spec declares them, and its rendered output is byte-identical to
before. The rewrite also drops prance, openapi-spec-validator and five more
transitive packages.
tokenizers stays at 0.22.2: transformers 5.8.1 caps it at <=0.23.0 and no
such release exists, so it moves with the transformers cap or not at all.
`uv lock --upgrade` plus the two code changes the new versions need.
firecrawl-anydoc 0.2.4 raises a dedicated `NeedsOcrError` where 0.2.3 raised
`UnsupportedError("... OCR is required")`, so the anydoc parser no longer
recognised a scanned PDF: the fallback still ran, but a near-empty result was
stored as an empty document instead of failing with the OCR_ENABLED hint.
`_needs_ocr` now accepts both spellings and looks the class up lazily, so an
older anydoc keeps working. 0.2.4 also refuses the CID-font NDA fixture
outright rather than dropping its Chinese column silently, so the PDF
trust-check tests stub that dropped output against the fixture's real bytes
(the check's own inputs) and a new test pins the refusal path.
ruff 0.16 widened its implicit default rule set, turning the dev-group bump
into 7131 findings across the tree. `.ruff.toml` now states the historical
selection (E4, E7, E9, F) explicitly and the CI pin moves to the locked
0.16.7, so lint no longer drifts with the version.
pip install docsgpt now brings the web UI with it: `docsgpt api` serves the
API and the UI on one port.
- scripts/build_frontend.sh builds the frontend into docsgpt/static
(gitignored) the way the frontend image does: .env.development as the
production baseline, and index.html loading /config.js ahead of the
bundle. hatch admits the directory into the wheel and the sdist through
`artifacts`; the package workflows run the script before `uv build` and
fail if the wheel lacks the UI. The backend image keeps ignoring it.
- docsgpt/ui.py serves the build in front of Flask: files as they are,
hashed assets immutable, Flask's own path prefixes (taken from its URL map,
so new blueprints need no registration) passed through, every other GET
rendered as index.html for the client-side router. /config.js is generated
per request with VITE_API_HOST and VITE_BASE_URL set to the page's origin,
VITE_* environment variables winning. SERVE_UI=false leaves the API alone.
- docsgpt api configures gunicorn in code (gunicorn.app.base.Application)
instead of rewriting sys.argv, so the SIGUSR2 re-exec that gunicorn uses
for zero-downtime upgrades runs the docsgpt console script again and
works; verified with a live handover.
- Docs: the pip page says the UI is included, that DOCSGPT_HOME and
DOCSGPT_ENV_FILE are process environment variables rather than .env
entries, and the settings page describes SERVE_UI.
- docsgpt api binds 127.0.0.1 by default, like gunicorn and uvicorn do;
--host 0.0.0.0 exposes it. The docs say so.
- The embedded Milvus and LanceDB defaults derive from the data home, so
they follow DOCSGPT_HOME like the faiss indexes and uploads do. A checkout
run from its root and the Docker image resolve to the same paths as before.
- The docs and the pyproject comment describe the CPU torch install as two
steps (torch and torchvision from the PyTorch CPU index first, then the
docling extra): pip picks the highest version across indexes, so
--extra-index-url only yields the CPU build while that index keeps pace
with PyPI.
- AGENTS.md separates DOCSGPT_HOME (moves the data home) from
DOCSGPT_ENV_FILE (selects the .env file); the docs example uses a password
placeholder.
pip install docsgpt (extras: docling, milvus) installs the backend with a
docsgpt command: api, worker, migrate, prefetch-models, verify-offline,
reembed. Second step of the PyPI work after the package rename.
- hatchling build; the version comes from docsgpt/version.py. The wheel is
the docsgpt package with the data it reads at runtime (prompts, model
catalogs, seed config, alembic.ini and migrations) and without the
Dockerfile, the exported requirements, the sample index and local runtime
data. The application import alias stays checkout-only. uv sync installs
the package editable now that [tool.uv] package = false is gone.
- docsgpt/cli.py: api (gunicorn + BoundedDrainUvicornWorker with the image's
flags, --reload for uvicorn), worker (Celery worker with beat embedded,
--no-beat/-Q/--concurrency/--pool, solo pool on macOS), migrate, and
argument pass-through to the maintenance scripts. --help imports no app.
- docsgpt/core/paths.py: runtime data lives in a data home (DOCSGPT_HOME,
else the checkout, else cwd); DOCSGPT_ENV_FILE overrides the env file.
Settings, the dotenv load, LocalStorage and the internal upload route use
it instead of "three directories above this file", which is site-packages
for an installed package. A checkout and the Docker image behave as before.
- [project] dependencies are compatible ranges so the package installs next
to other packages; uv.lock resolves to the same versions and the exported
requirements files are unchanged.
- package-build.yml builds and checks the wheel on PRs and installs it into
a clean venv; pypi-publish.yml publishes on a published release through
trusted publishing (environment pypi), or to TestPyPI on a manual run.
- Docs: Deploying -> Install with pip. AGENTS.md notes the package.
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
requirements.txt pinned torch and transformers in core although only docling
needs them, and on Linux torch pulls the CUDA 13 stack: 2.7 GB of the 3.0 GB
wheel download. Direct dependencies now live in pyproject.toml, uv.lock pins
everything, and application/requirements*.txt are exported from the lock by
scripts/export_requirements.sh (each file is the core set plus one extra).
The docling extra pins torch/torchvision/transformers itself and, on Linux,
resolves torch from the CPU-only PyTorch index (no nvidia packages). milvus
(pymilvus + milvus-lite, which pulls pyarrow) is the second extra.
application/core/optional_deps.py is the one place install hints come from;
the milvus store and the docling call sites use it so a missing extra fails
with the exact command to run.