The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
Review follow-ups.
A BOM told the sniff which encoding to read, but was also taken as the
verdict: three prepended bytes let any binary through, including as
notes.txt. A BOM now only selects the test — UTF-8 falls through to the
byte rules on the remainder, UTF-16/32 decode and judge the characters
(NUL, unprintable, or replacement chars from bytes the decoder could not
read). Real Notepad-Unicode text still passes, mp4-behind-a-BOM does not,
in either language.
The gate treated the full parser table as a given, but without docling the
fallback extractor has no .tif/.tiff/.bmp/.webp/.vtt/.xml handler, so those
suffixes skipped the content check and reached the plain-text fallthrough —
the original bug, one install away. The worker now passes the keys of the
extractor it actually built, making the second gate stricter than the
route's static one rather than a copy of it.
Cache bins: extraction coerced a missing value to 0 and only non-zero bins
were recorded, so a provider reporting cached_tokens=0 persisted as NULL —
indistinguishable from "not reported", and OpenAI reports exactly that on
every uncached request. Bins are now carried as Optional and recorded when
not None, which is what the nullable columns and the NULL-means-unknown
comment already assumed. Anthropic's cache_read/cache_creation bins had the
same shape and are fixed alongside; the int-or-None coercion is shared in
llm/base.py.
Review follow-ups on the attachment gate.
.txt was listed as parser-backed, but it has no parser — it *is* the
plain-text fallthrough. That let it skip the content check, so renaming a
video to notes.txt walked straight back into the bug the gate exists for
(verified: 5132 chars of binary "extracted" and stored). The list is now
exactly the file extractor's keys, .txt included in the content check like
any other unparsed suffix, and the drift test asserts equality rather than
containment. The sniff now recognises a UTF-16/32 BOM as text, so a
Notepad "Unicode" .txt is not caught by the NUL-byte rule.
The picker's accept filter listed parser-backed suffixes only, hiding .txt,
.py and .log — files the gate reads happily — behind "All files". It now
carries text/* as well, so it can never be narrower than what the upload
accepts.
A rejected batch carries one errors entry per file, but the non-200 branch
applied the top-level message to every chip, so two files failing for two
reasons both reported the first one. Reasons are now matched by
upload_index, with the top-level message as fallback.
_get_store_attachment_user_error no longer reads str(exc): the
unsupported-type message is rebuilt from the filename, so no exception
state can reach a response body (CodeQL py/stack-trace-exposure).
A chat attachment with no parser fell through to SimpleDirectoryReader's
plain-text open(), so a phone-uploaded video was "extracted" into megabytes
of binary garbage, truncated, and stored with extraction.status == "ok".
Gate attachments in two tiers instead. A suffix with a dedicated parser is
admitted on its name — a PDF is binary and parses fine. Anything else has to
read as text: the first 8KB are sampled and refused on a NUL byte or too many
other control bytes. That keeps source, config and log files working through
the plain-text fallthrough, and keeps out videos, archives and renamed
binaries alike. The route checks the staged spool before anything is stored
or queued; the worker repeats the check where the local file exists, raising
the non-retryable AttachmentRejectedError.
SUPPORTED_ATTACHMENT_EXTENSIONS gains the parser-backed suffixes it was
missing (.tiff, .tif, .bmp, .webp, .vtt, .xml) and is now exactly the file
extractor's keys plus .txt, with a test asserting the two agree. The composer
applies the same rule client-side, so an unsupported file is named before it
costs an upload, and a test pins the frontend list to the backend one.
Attachment failures now show their reason inline under the chips rather than
only in a hover tooltip, which a touch user can never see, and only after a
send was attempted. Dropped `accept` from the dropzone: it discarded rejected
drops with no feedback and disagreed with the server about text files.