Exposes the three per-source graph options in the source's retrieval settings,
shown only when the retriever is graphrag: where the walk starts (entities or
relationships), whether passages join the walk, and whether vector hits are
blended in. Defaults match the backend's measured-best configuration and are
filled in for sources saved before the options existed. A note points at the
"search tool" exposure, which is what offers the graph to an agent.
- Ship tiktoken's cl100k_base inside the package and build the encoding
from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
temp dir, and read tokenizer.json and repo metadata from that cache, so
a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
the endpoints return 404, audio files fail to ingest with a clear
message, /api/config reports tts_available/stt_available, and the UI
hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
The connector OAuth popup posted the session token to window.opener with
a '*' target origin, so any page that opened the popup received it. An
attacker with an account on a multi-user deployment could start a flow for
their own pending session, get a victim to finish the provider consent, and
receive a token backed by the victim's Drive/SharePoint/Confluence tokens.
- Post popup results only to allowed frontend origins: the callback origin,
OIDC_FRONTEND_URL, the new CONNECTOR_ALLOWED_ORIGINS, and localhost:5173
when the callback runs on a loopback host.
- Render the success page from the callback itself so the token never
appears in a URL; callback-status ignores session_token/user_email params.
- ConnectorAuth accepts messages only from the popup it opened, on the
callback origin reported by /api/connectors/auth.
- /api/connectors/disconnect requires auth and only deletes the caller's
session.
- /api/connectors/sync and /api/remote reject session tokens the caller
does not own.
Fixes#2766
The button is disabled once isCopied is true, but copy() is now async, so
isCopied and the disabled prop only catch up after it resolves. Clicks
landing inside that window both passed the guard and started a write.
Nothing leaked, since the existing clearTimeout already handled the
duplicate timer, but the window did not exist before copy-to-clipboard 4
made the call async. A ref cleared in finally closes it.
Raised by CodeRabbit on #2762. 598 tests pass, build and lint unchanged.
eslint.config.js imports @eslint/js and globals but neither was declared;
both resolved transitively through eslint. Declaring them fixes a latent
break and is a prerequisite for ESLint 10, which stops providing them.
ESLint 10 itself is not taken. eslint-plugin-react 7.37.5 is the latest
release and peers at eslint "^3 || ... || ^9.7", and under 10 every rule
load throws "contextOrFilename.getFilename is not a function" because
ESLint 10 removed context.getFilename(). The frontend stays on eslint 9
until that ships; note npm already marks 9.x as no longer supported.
eslint-plugin-n goes to 18, though the flat config does not reference it
(nor eslint-plugin-import, eslint-plugin-promise or eslint-config-prettier
-- all four are unused devDependencies).
Lint output unchanged from baseline.
mermaid 12 needs no source change; mermaidSecurity.ts and its strict-mode
init-directive test pass unchanged.
Adds a lodash-es ^4.18.1 override. mermaid 12 depends on chevrotain ~11.1.2
(mermaid 11 resolved chevrotain 12, which has no lodash-es), and chevrotain
11 pulls lodash-es <= 4.17.23, carrying GHSA-r5fr-rjxr-66jc and
GHSA-f23m-r3pf-42rh, both high. 4.18.1 is the fixed line. Without the
override the frontend gains 5 high-severity advisories.
Verified rendering in Chromium rather than the unit suite, since happy-dom
has no SVG getBBox and cannot render any mermaid version. Flowchart,
sequence, pie, class, state, gantt, ER and mindmap all produce SVG on 12,
matching 11 to the byte except classDiagram (+17 bytes). A strict-mode
`click ... "javascript:..."` target is still stripped.
598 tests pass, build clean, 0 advisories.
react-markdown 10 removed the `className` prop and no longer renders a
wrapper element; passing className now throws at render time. Each of the
three call sites gets an explicit wrapping div carrying the same classes,
which reproduces the v9 DOM exactly.
The large diff is re-indentation from the added wrapper. `git diff -w`
shows the only substantive change per file is the moved className.
Fixes the 19 tests that failed on the upgrade. 598 tests pass, build clean,
lint unchanged.
copy-to-clipboard 4 returns Promise<boolean> instead of boolean, so
CopyButton's handleCopy awaits it. Without the await the success branch
always ran, since a Promise is always truthy. Awaiting also routes a
rejected copy into the existing catch.
vite-plugin-svgr 5 and @shadcn/react 0.3.1 needed no source changes.
598 tests pass, build clean, lint unchanged.
The chat selection was inferred by diffing the refreshed source list
against the snapshot taken before the upload, so any source that appeared
meanwhile -- another upload finishing, a new team share -- could be picked
instead. The upload response already carries the id; match on it, and
select nothing when it is absent rather than guessing.
Drops the now-unused sourceDocs read and its stale dependency.
selectedSourceIdsFromAgent preferred the extra sources over the legacy
single source instead of merging them. Since every save writes both
fields, opening an agent that held both and saving -- even just a rename
-- silently detached the primary source. Retrieval reads both, so merge
them.
The edit form also reported unsaved changes the moment it loaded: the
snapshot taken after the fetch defaulted the retriever to "classic" and
left the default model empty, while the effects that run straight after
rewrite both. Normalise the snapshot through the same serializer and
rules the effects apply.
A source attached to a team-shared agent that the caller cannot list --
the owner's private source -- rendered no picker row, so it showed as
attached with nothing to switch it off. Append rows for selected ids
missing from the list, and memoise the label lookup the trigger and the
picker now share.
Uploading from the agent builder repointed the open conversation onto the
new source and persisted it, because the upload modal always dispatched
the chat selection. Gate that behind a prop the builder opts out of. The
wiki source is created synchronously and never refreshed the source list,
so the builder selected an id the picker could not render and the trigger
read "External KB"; refresh on every completion path that skips ingest
tracking.
Removing the synthetic "Default" source made an empty source list
reachable, which the stored-selection reconciler skipped, leaving a
deleted source checked forever with no way to clear it. Treat an empty
list as a real answer and only skip when the list has not loaded.
Finally, finish the picker extraction: the chat input still hand-rolled
the item mapping, so team users saw the same sources grouped in the agent
builder and flat in chat. Point it at the shared helper, with its own
locale keys, and make the sources link a router navigation instead of a
full page reload.
Review follow-ups:
- The agents list and the pinned list dropped agents that had neither a
source nor a retriever. Publishing such an agent is now allowed, so it
vanished from the list right after publish. Both filters are gone; a
source-less agent skips retrieval and is still runnable. Tests cover
both listings.
- The agent form selected an uploaded source by diffing the source
catalog before and after the upload, which would also pick up any
source created or shared in the meantime. Upload now passes the id of
the source it created to onSuccessfulUpload, and the form selects only
that id.
The sources list used to start with a fake "Default" entry that had no id
and, at run time, meant "no source, skip retrieval". The agent form
pre-selected it, snapped back to it when the last source was deselected,
and refused to publish without it, so a new agent always looked like it
had a knowledge base when it had none.
Backend
- /api/sources returns only ingested sources; no placeholder row.
- Publishing an agent no longer requires a source on create or update.
The legacy "default" value is still accepted and maps to NULL.
Frontend
- The agent source picker starts empty, can be cleared, and shows a hint
that a source-less agent answers from the model and its tools only.
- The picker groups sources into "Your sources" and "Shared with team"
when any team-shared source exists, shows "N sources selected" for a
multi-selection, and gets the same "Go to Sources" / "Upload new"
footer as the chat picker. A source uploaded from the form is selected
when it lands.
- Source selection serialisation and the picker id live in one helper
shared with the chat picker; the four copies in the form are gone.
- The client no longer seeds a placeholder source in the store, and the
dead auto-select of a "default" document is removed.
- The frontend image ran the Vite dev server in development mode, so
.env.development supplied its defaults (notification banner, Google client
id, local API host). The static build only loads .env.production, so the
build stage now copies .env.development in as the baseline and the compose
files pass every VITE_* the app reads through from .env; the runtime script
skips empty values so a blank passthrough keeps the build-time default.
.dockerignore kept only the .local variants out.
- VITE_DISABLE_SOURCE_FE disables sources only when it is the string true.
- DoclingParser: find_spec raises when docling itself is absent; the install
hint now covers that path, with a regression test.
- verify_offline: direct tests for verify(); the PR image check builds and
verifies the -docling variant as well as slim.
- Workflows this branch adds or rewrites pin actions by commit, pass the
release tag through env instead of template expansion, and do not persist
checkout credentials.
- OCR guide no longer claims pre-built images never include docling.
Backend (arc53/docsgpt): 4.5 GB compressed -> 0.9 GB with both embedding
models and tiktoken baked in.
- torch/transformers gone from the default install (docling extra only).
- Ubuntu 24.04 ships python3.12: no deadsnakes PPA, no software-properties-
common; every pin is a wheel, so no gcc/g++/rust in the builder.
- COPY --chown and a prefetch that runs as the process user replace the
trailing chown -R, which duplicated the 600 MB model layer.
- .dockerignore keeps __pycache__, .coverage, local indexes and .env out.
- EXTRAS build arg (INSTALL_DOCLING kept as an alias); the docling variant
also bakes docling's layout/table/RapidOCR models (DOCLING_ARTIFACTS_PATH)
and tesseract, and drops only the discovery documents of Google APIs the
app never builds.
- FLASK_DEBUG env removed (unused); OCI labels added.
Frontend (arc53/docsgpt-fe): 302 MB Vite dev server -> 25 MB static build
behind nginx. VITE_* variables are injected at container start into
/config.js and read through src/env.ts, so the image no longer needs a
rebuild per deployment; docker-compose.yaml keeps hot reload via the dev
target.
Publishing: every release and develop build now pushes a slim tag and a
-docling tag (docling engine + models + tesseract). docker-compose-hub.yaml
takes DOCSGPT_IMAGE_TAG / DOCSGPT_IMAGE_VARIANT; docker-compose-standalone.yaml
runs the stack from pre-built images without a checkout and is attached to
each release. setup.sh selects the -docling variant for OCR instead of
requiring a local build. A new workflow builds the image on PRs that touch
it and runs verify_offline under --network none; lint checks the exported
requirements match uv.lock.
Conflicts, and how each was taken:
- application/core/settings.py — ours. The renamed OCR_ENABLED /
OCR_ATTACHMENTS_ENABLED / OCR_MIN_CHARS_PER_PAGE accept main's
DOCLING_OCR_* spellings as AliasChoices, so nothing is dropped.
- application/Dockerfile — both. Main's install layers plus the
INSTALL_DOCLING build arg.
- application/parser/file/constants.py — both imports.
- deployment/docker-compose.yaml — both. The INSTALL_DOCLING /
INSTALL_TESSERACT build args on backend and worker, and main's
-Q docsgpt,parsing,embeddings, which query embedding needs.
- tests/conftest.py — theirs. Both sides fixed the same pytest-postgresql
9.0.0 autocommit= breakage; main's spelling is the one already on main.
- application/requirements.txt — the comments claimed different reasons
torch is in core. Main's is the true one now: it removed
sentence-transformers, so docling is torch's only remaining consumer.
Two things the merge broke without conflicting:
- onnxruntime. This branch moved it out of core into the docling extra;
main meanwhile made it the runtime local embeddings execute on
(fastembed). Git took the deletion, leaving fastembed with no pinned
runtime in a repo that pins everything. Restored to core, and no longer
pinned twice from the extra.
- The frontend copy of ATTACHMENT_PARSER_EXTENSIONS. The backend list is
derived and picked up the anydoc suffixes; the hand-kept frontend mirror
did not, so the composer would refuse files the API accepts.
tests/parser/file/test_constants.py is what caught it.
ruff, pytest (9897 passed), frontend build and docs build all pass. The
image build is unverified: no Docker daemon on this machine.
Review follow-ups.
A BOM told the sniff which encoding to read, but was also taken as the
verdict: three prepended bytes let any binary through, including as
notes.txt. A BOM now only selects the test — UTF-8 falls through to the
byte rules on the remainder, UTF-16/32 decode and judge the characters
(NUL, unprintable, or replacement chars from bytes the decoder could not
read). Real Notepad-Unicode text still passes, mp4-behind-a-BOM does not,
in either language.
The gate treated the full parser table as a given, but without docling the
fallback extractor has no .tif/.tiff/.bmp/.webp/.vtt/.xml handler, so those
suffixes skipped the content check and reached the plain-text fallthrough —
the original bug, one install away. The worker now passes the keys of the
extractor it actually built, making the second gate stricter than the
route's static one rather than a copy of it.
Cache bins: extraction coerced a missing value to 0 and only non-zero bins
were recorded, so a provider reporting cached_tokens=0 persisted as NULL —
indistinguishable from "not reported", and OpenAI reports exactly that on
every uncached request. Bins are now carried as Optional and recorded when
not None, which is what the nullable columns and the NULL-means-unknown
comment already assumed. Anthropic's cache_read/cache_creation bins had the
same shape and are fixed alongside; the int-or-None coercion is shared in
llm/base.py.
Review follow-ups on the attachment gate.
.txt was listed as parser-backed, but it has no parser — it *is* the
plain-text fallthrough. That let it skip the content check, so renaming a
video to notes.txt walked straight back into the bug the gate exists for
(verified: 5132 chars of binary "extracted" and stored). The list is now
exactly the file extractor's keys, .txt included in the content check like
any other unparsed suffix, and the drift test asserts equality rather than
containment. The sniff now recognises a UTF-16/32 BOM as text, so a
Notepad "Unicode" .txt is not caught by the NUL-byte rule.
The picker's accept filter listed parser-backed suffixes only, hiding .txt,
.py and .log — files the gate reads happily — behind "All files". It now
carries text/* as well, so it can never be narrower than what the upload
accepts.
A rejected batch carries one errors entry per file, but the non-200 branch
applied the top-level message to every chip, so two files failing for two
reasons both reported the first one. Reasons are now matched by
upload_index, with the top-level message as fallback.
_get_store_attachment_user_error no longer reads str(exc): the
unsupported-type message is rebuilt from the filename, so no exception
state can reach a response body (CodeQL py/stack-trace-exposure).
A chat attachment with no parser fell through to SimpleDirectoryReader's
plain-text open(), so a phone-uploaded video was "extracted" into megabytes
of binary garbage, truncated, and stored with extraction.status == "ok".
Gate attachments in two tiers instead. A suffix with a dedicated parser is
admitted on its name — a PDF is binary and parses fine. Anything else has to
read as text: the first 8KB are sampled and refused on a NUL byte or too many
other control bytes. That keeps source, config and log files working through
the plain-text fallthrough, and keeps out videos, archives and renamed
binaries alike. The route checks the staged spool before anything is stored
or queued; the worker repeats the check where the local file exists, raising
the non-retryable AttachmentRejectedError.
SUPPORTED_ATTACHMENT_EXTENSIONS gains the parser-backed suffixes it was
missing (.tiff, .tif, .bmp, .webp, .vtt, .xml) and is now exactly the file
extractor's keys plus .txt, with a test asserting the two agree. The composer
applies the same rule client-side, so an unsupported file is named before it
costs an upload, and a test pins the frontend list to the backend one.
Attachment failures now show their reason inline under the chips rather than
only in a hover tooltip, which a touch user can never see, and only after a
send was attempted. Dropped `accept` from the dropzone: it discarded rejected
drops with no feedback and disagreed with the server about text files.