Graph retrieval tied plain vector search at best and never beat it. Measured
across five corpora, the bottleneck was seeding, not the graph: the walk
started from nodes whose embeddings were computed from bare entity names, and
a whole question shares almost nothing with a name like "Quill".
Extraction now embeds each node from "name (type): description" and each
relationship as the fact it asserts ("Alder streams_to Quill: ..."), stored on
a new nullable graph_edges.fact_embedding column that ensure_vector_schema adds
in place. Entity names are canonicalised (case, punctuation, word breaks and a
cautious plural) so "VECTOR_STORE" and "vector stores" land on one node. Extraction calls run
concurrently (GRAPHRAG_EXTRACTION_WORKERS, default 8) while embedding and graph
writes stay serial on the task thread, so ordering and idempotency are
unchanged; that measured 8.4x faster with identical output.
Retrieval gains per-source options, stored under retrieval.graph and read live
at query time:
- seed_strategy: start from matching entities (default) or matching
relationships, which can reach an entity the question never names;
- passage_nodes (on): walk the source's passages alongside entities, with
PageRank damping 0.5 instead of 0.85;
- blend_vector (on): fuse the graph ranking with the source's vector ranking
by reciprocal rank.
The defaults are the measured-best configuration. Through GraphRAGRetriever,
the new seeding moved recall@4 from 0.41 to 0.68 on a multi-hop corpus and
from 0.50 to 1.00 on the docs corpus, and regressed none of the corpora
measured. Existing graphs keep name-only embeddings until rebuilt.
The docs site failed to build: a bare "<= 1" in the prose of the
generated page is parsed by MDX as the start of a JSX tag ("Unexpected
character '=' before name"). Constraints are rendered as code spans now,
where MDX leaves them alone, and a test rejects any bare <, { or } outside
a code span so a future description cannot reintroduce the failure.
Verified with a local next build of the docs site.
Review follow-up. The per-group secret validators normalised a hand-picked
list of API keys, which left other optional credentials and overrides
(OPEN_ROUTER_API_KEY, S3 and Daytona keys, ELASTIC_PASSWORD, the OIDC
trio, connector client ids, MICROSOFT_AUTHORITY, MCP_OAUTH_REDIRECT_URI)
holding the literal "None" or "" a .env file spells "unset" with, so
truthiness checks and fallbacks downstream saw a value. One rule on the
group base replaces those lists: every Optional[str] field maps "", "None"
and whitespace to None and strips real values. Plain str fields are left
alone. The OIDC required-settings check therefore also rejects those
spellings.
EMBEDDINGS_POOLING is Literal["cls", "mean"] with case-insensitive
parsing; its consumer silently ignored anything else.
Bounds added where the consumer rejects or misbehaves on the value:
SCHEDULE_RUN_OUTPUT_RETENTION_DAYS and MESSAGE_EVENTS_RETENTION_DAYS (the
cleanup repositories raise on <= 0), EMBEDDINGS_DELEGATE_TIMEOUT, the
remote-device idle/pairing/invocation TTLs and CELERY_VISIBILITY_TIMEOUT
(> 0), REMOTE_DEVICE_CMD_QUEUE_TTL_SECONDS (> 605, the documented drain
deadline), GRAPHRAG_MAX_CHUNKS_FOR_EXTRACTION (>= 0; negative would slice
the pending list from the end).
The generated reference now renders generic type arguments
(dict[str, int] rather than dict).
`docsgpt up --native` installs services meant to outlive the shell. Development
wants the opposite, and until now it meant three terminals from the guide:
uvicorn, celery, and vite.
`docsgpt dev` runs this checkout's API and worker as children of one terminal,
both restarting when a file is saved, their output interleaved and labelled, and
Ctrl-C stopping them together. `--ui` adds the Vite dev server, `--mock-llm`
runs the bundled mock model so no API key is needed, and `--no-worker` leaves
the worker to your editor's debugger. Celery has no reloader of its own, so the
worker is wrapped in watchfiles when it is installed, and runs plain when it is
not.
Alongside it, the commands a dev loop keeps reaching for:
- `docsgpt doctor` checks what usually breaks a new setup: PostgreSQL answering
and its schema matching this version, Redis answering, a model provider being
configured, and the port being free.
- `docsgpt restart [api|worker]` bounces services without rewriting settings or
rerunning migrations, which `down` plus `up` did.
- `docsgpt logs -f` follows a native install instead of telling you to run
`tail -f` yourself.
- `docsgpt env set` applies itself to a running native install rather than
asking you to run `docsgpt up` again to change one value.
Two bugs found on the way, both older than this change:
- `docsgpt api --reload` watched the working directory, which in a checkout is
178,425 files: .venv, node_modules, and the indexes/ and inputs/ the app
writes to while ingesting, so the server restarted itself mid-request. It
watches the package now — 1,217 files.
- The VS Code "Flask Debugger" ran `flask run`, which serves only the WSGI app:
/mcp, the SSE streams and artifact downloads 404 under it. The guide warned
about this in prose while the debug config did it anyway. It runs uvicorn on
the ASGI app now, like production.
About 85 call sites read a setting as getattr(settings, "NAME", fallback),
each carrying its own copy of the default. Every one of those names is a
field with a default on the model, so the fallback could never apply to
the real settings object; it only masked drift. Two had drifted:
- OPENAI_PROMPT_CACHE_KEY defaults to True on the model but the reader
fell back to False, and two test stubs relied on that.
- SharePoint's MICROSOFT_AUTHORITY fallback to
https://login.microsoftonline.com/<tenant> never fired, because the
attribute always exists (as None), so MSAL got authority=None. The
connector now derives the tenant authority when the setting is unset,
as its test always assumed.
Four places read EMBEDDINGS_KEY straight from os.environ, skipping the
"None"/"" normalisation the model applies; they read the setting now.
Test stubs that replaced a module's settings with a SimpleNamespace list
every setting the code under test reads.
SAGEMAKER_REGION, SAGEMAKER_ACCESS_KEY and SAGEMAKER_SECRET_KEY survive
only as a fallback for the S3_* credentials. They carry
Field(deprecated=...) now, so any read emits a DeprecationWarning naming
the replacement and the generated reference shows the notice. The S3
store is the one sanctioned reader; it silences that warning locally
because it already logs its own operator-facing one when the fallback
is actually used.
DEFAULT_MAX_HISTORY was referenced nowhere. RETRIEVERS_ENABLED was read by
no code at all, while two docs pages described it as an enforced
allow-list; both the setting and those claims are removed.
The "AUTH_TYPE=oidc requires OIDC_ISSUER, OIDC_CLIENT_ID and
OIDC_FRONTEND_URL" check lived in app.py, so it only ran when the Flask
app was imported; a worker or script with the same misconfiguration
started fine. It is now a model validator on the auth group and runs
wherever Settings is loaded, with the same message.
DEPLOYMENT_TYPE, which app.py read straight from the environment to
decide whether a missing JWT_SECRET_KEY is fatal, is a documented
setting on the server group now, so it shows up in the reference like
every other variable the app reads.
Enum-like settings whose allowed values were only listed in a comment are
now Literal types, so a typo fails at startup with a message naming the
allowed values instead of falling through to a default with a warning
(or, for VECTOR_STORE, failing on first use):
AUTH_TYPE, VECTOR_STORE, STORAGE_TYPE, URL_STRATEGY, OCR_BACKEND,
OCR_ENGINE, SANDBOX_BACKEND, DOC_PARSER_ENGINE, TTS_PROVIDER, STT_PROVIDER
Each keeps a before-validator that strips and lower-cases the value, since
the registries that consume them already lower-cased at the use site, and
AUTH_TYPE maps the "None"/"none"/"" spellings a .env file carries to None
(it was the string "None" before, which only worked because nothing
compared against it). An empty TTS/STT provider still means "off".
LLM_PROVIDER stays a plain str because providers are plugin-extensible.
Containers are typed (dict[str, int], list[str], dict[str, Any]) instead
of bare dict/list, six fields that were Optional with a non-None default
are plain, and integer settings whose description already states a range
carry it as a constraint (ge=0 for "0 disables", ge=1 for counts that
cannot be zero, 0 < threshold <= 1).
The hand-maintained settings page documented 95 of 258 settings and
.env-template 42, and both drifted as fields were added. The field
descriptions now live on the model, so the reference is rendered from it:
python -m docsgpt.core.settings.reference --write
writes docs/content/Deploying/Settings-Reference.mdx, one section per
settings group with each field's type, default, constraints, aliases and
description. --check reports a stale page, and tests/core/test_settings.py
fails when the checked-in page no longer matches the definitions, so a
new setting cannot land undocumented.
test_settings.py also pins the composition contract: every group field is
a flat Settings attribute, no field is defined twice, every field has a
description, and the secret-normalising validator of every group is
applied (the case that a shared method name would silently drop).
The App Configuration page points at the reference instead of at
settings.py, and the reference is listed in the Deploying navigation.
docsgpt/core/settings.py had grown to 258 fields in one 600-line class,
touched by about two commits a week, with related settings scattered
(GitHub ingest caps inside the embeddings block, API keys in four places,
the OpenAI Responses knobs 100 lines from the other OpenAI fields).
It is now a package: one module per domain (auth, llm, embeddings,
retrieval, vectorstores, database, workers, ingestion, ocr, storage,
connectors, server, events, agents, guardrails, scheduler, sandbox,
speech), each a SettingsGroup owning its fields and validators, composed
by multiple inheritance into the same flat Settings class. Every
attribute name, type, default, alias and constraint is unchanged, so
settings.NAME reads, .env files and test monkeypatches all keep working;
the import path docsgpt.core.settings is the package. Settings.normalize_api_key
is kept as a classmethod for callers that reuse it.
The comment above or beside each field became its Field(description=...),
so the definitions are visible to tooling; the next commit generates the
docs reference from them.
Pitfall recorded for future groups: pydantic collects validators by
method name across the MRO, so two groups naming a validator the same
would silently keep only one. Each group's validator has a unique name.
From review of #2800:
- systemd `enable --now` starts nothing when the unit is already active, so a
second `up --native` kept the old ExecStart and left the API on its previous
port. start enables and then restarts, as the launchd path already did by
booting the job out first.
- An explicit `home` now wins over XDG_CONFIG_HOME, which is what callers pass
it for.
- The ExecStart program must be a real executable: when `docsgpt` is not on
PATH, sys.argv[0] is accepted only if it can be run, and otherwise the
failure is raised before any unit is written.
- `up --native` over a directory holding a Docker install now refuses and says
how to proceed, instead of starting native services beside containers that
down, status and uninstall would no longer see.
- Docs: without a terminal only --postgres-uri is required, and the Windows
fallback names `docsgpt beat`, which the worker cannot embed there.
SystemdServices was the least covered part of the module and cannot be run on
this machine, so it now has tests for install, start, stop, remove, is_running
and a failing systemctl.
From review of #2799:
- The archive is created 0600 rather than at the process umask: it holds the
install's data, and --with-settings puts .env and its secrets in it.
- restore validates everything the manifest declares before the stack is
stopped, so a damaged archive fails while DocsGPT is still running rather
than after `compose down` has taken it away.
- Only the volumes a backup is made of are restored. A hand-made manifest can
no longer point import_volume at postgres_data, whose contents it empties.
- psql runs with ON_ERROR_STOP=on, so a restore that fails halfway cannot
start DocsGPT again and call it a success.
- import_volume unpacks into the container's own filesystem first and clears
the live volume only once the tar has come out whole, so a corrupt one
leaves the volume as it was.
- The backend and the worker stop while the archive is made and start again
even if the dump fails, so the dump and the volume tars describe the same
moment instead of drifting apart as ingestion writes.
Run DocsGPT without Docker: the API and the worker each become a service
on the machine itself, a launchd agent on macOS and a systemd user unit
on Linux, pointed at a PostgreSQL and a Redis that already run.
`docsgpt up --native --postgres-uri ... --redis-url ...` writes the same
.env a Docker install uses, applies the migrations and starts both
services. status, logs, down and uninstall work on a native install the
same way they do on a Docker one, and never touch the database or Redis:
they were the user's to begin with.
One Redis URL covers the broker, the result backend and the cache on
three consecutive databases, starting at the one the URL names, so a
Redis that already holds something else can be shared.
Windows has neither service manager, so native mode refuses it and says
what to do instead.
`docsgpt backup` writes one archive holding a pg_dump of the database, a tar
of each data volume and a manifest of what it came from; `docsgpt restore`
puts it back over an install. The settings file is left out unless
--with-settings asks for it, since it holds the install's secrets, and a
backup taken with a newer DocsGPT is refused without --force.
The volume tars go through the image the install already runs, so a backup
pulls nothing extra, and compose calls can now redirect stdout and stdin so
the dump never passes through this process.
- install.sh saves the get.docker.com and uv installers to a file and runs
them only after the download finished, so a cut-off transfer runs nothing.
- Neither installer prints DOCSGPT_PACKAGE, which may be a URL with
credentials.
- The CI step assigns the wheel path before exporting it, so a missing wheel
fails instead of installing from PyPI.
- Docker-Deploying shows one code block per platform; Quickstart names the
/opt/docsgpt home used for root on Linux.
deployment/install.sh (curl | bash) and install.ps1 (irm | iex) check for
Docker, install uv when it is missing or older than 0.8 (pinned 0.12.15 via
Astral's installer), install or upgrade the docsgpt package with
`uv tool install`, and hand the terminal to `docsgpt up` with any arguments.
On Linux without Docker the shell installer offers get.docker.com. Both run
entirely inside a function, so a download cut short runs nothing.
Releases attach both scripts next to the Compose file, which is where
docs.ac/install and docs.ac/install.ps1 will point. installer-lint.yml runs
shellcheck and the PowerShell parser; docker-image-verify.yml now installs
through install.sh. README, Quickstart, Docker-Deploying and the changelog
lead with the one-liner.
- wait_healthy starts no request once the deadline is reached, and neither
its pauses nor the requests after the first run past it; the first attempt
still always runs (status uses a zero timeout).
- envfile writes $ as $$ inside double quotes, which Compose interpolates,
and reads $$ back as $, so a value such as pa$w'rd reaches the container
unchanged.
- Upgrading no longer describes the working directory as the data home.
Docker-Deploying gains a `docsgpt up` section, Pip-Install and Upgrading
describe the ~/.docsgpt/server data home, and the changelog covers both.
docker-image-verify.yml installs the wheel and runs `docsgpt up`, `status`,
a second `up` that must keep the secrets, and `uninstall --purge` against
the image it built. The standalone Compose file maps host.docker.internal
to the host gateway, so a model server on a Linux host is reachable the way
`docsgpt up` suggests.
The backend image builds the web UI with scripts/build_frontend.sh and
serves it through docsgpt/ui.py, so the standalone Compose file drops the
frontend container. UI and API share port 7091, published on 127.0.0.1
unless DOCSGPT_BIND says otherwise. POSTGRES_PASSWORD is configurable, and
an optional https profile puts Caddy in front of a public domain.
docker-image-verify.yml starts the standalone stack on the image it built
and checks the API, the UI, /config.js and a client-side route on one port.
- Ship tiktoken's cl100k_base inside the package and build the encoding
from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
temp dir, and read tokenizer.json and repo metadata from that cache, so
a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
the endpoints return 404, audio files fail to ingest with a clear
message, /api/config reports tts_available/stt_available, and the UI
hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
The connector OAuth popup posted the session token to window.opener with
a '*' target origin, so any page that opened the popup received it. An
attacker with an account on a multi-user deployment could start a flow for
their own pending session, get a victim to finish the provider consent, and
receive a token backed by the victim's Drive/SharePoint/Confluence tokens.
- Post popup results only to allowed frontend origins: the callback origin,
OIDC_FRONTEND_URL, the new CONNECTOR_ALLOWED_ORIGINS, and localhost:5173
when the callback runs on a loopback host.
- Render the success page from the callback itself so the token never
appears in a URL; callback-status ignores session_token/user_email params.
- ConnectorAuth accepts messages only from the popup it opened, on the
callback origin reported by /api/connectors/auth.
- /api/connectors/disconnect requires auth and only deletes the caller's
session.
- /api/connectors/sync and /api/remote reject session tokens the caller
does not own.
Fixes#2766
A deflate-compressed TIFF under 1 MB can declare 144 million pixels and
take 1.2 GB to convert to PNG, and Pillow only warns below 179 million.
Read the dimensions from the header and refuse images over 40 million
pixels before any pixel data is decoded. Pillow's DecompressionBombError
is now raised as DocumentParseError, so the upload fails once instead of
being retried.
- Per-user SSE cap uses per-connection leases in a sorted set instead of
a shared INCR/DECR counter with a TTL. A stream that outlived the TTL
could let the counter expire, then decrement another stream's slot or
drive the count negative. Leases refresh while a stream sends frames
and age out when a stream dies without cleanup.
- Shielded cleanup is bounded per step: stream close and on_close in
ClosingStreamingResponse, unsubscribe and close in AsyncTopic, lease
release, and the replay-budget check, so a dead Redis connection can't
hold a request or a graceful shutdown.
- Artifact downloads send an ASCII filename plus an RFC 5987 filename*
for non-ASCII names, build headers before opening the file, and close
disk handles off the event loop.
- Type hints and docstrings on the new helpers.
Move GET /api/events, the remote-device command stream and artifact
downloads from Flask to Starlette routes mounted ahead of the Flask
catch-all. On Flask each held an a2wsgi threadpool slot for as long as
its response stayed open, and because uvicorn drops writes after a
client disconnects, a closed tab never released it.
- asgi_auth: one JWT/OIDC gate for Starlette routes; the chat reconnect
reader uses it too
- ClosingStreamingResponse closes the body iterator and releases the
SSE slot or file handle even when the client leaves before the first
frame
- AsyncTopic liveness probe replaces the sync client's socket_timeout
guard against half-open pub/sub sockets
- ASYNC_REDIS_MAX_CONNECTIONS sizes the async Redis pool; every open
stream holds a connection and redis-py defaults to 100
The changelog page has been an empty stub carrying only a title, and was
hidden from the sidebar. Fill it in from the 38 pull requests merged since the
0.19.0 tag, grouped by what a reader would look for, and link the GitHub
release notes for the full per-PR list and the upgrade guide for the steps an
existing deployment has to take. Earlier releases are not backfilled.
Now that the page has content, show it in the sidebar.
Next 16.3.5 and React 19.3 in docs/, plus the transitive refresh that came
with them (mermaid 11.17, katex 0.16.47, shiki 3.23, styled-components 6.5).
Two overrides were needed:
- zod ~4.3.6. nextra 4.6.1 declares `zod: ^4.1.12` but its MDX prop
validation breaks on zod >= 4.4.0, failing prerender for every page with
"Invalid input: expected nonoptional, received undefined -> at children".
Bisected: 4.3.6 builds, 4.4.0 through 4.6.2 all fail. nextra 4.6.1 is the
latest release, so there is no upstream fix to take yet. Drop this pin
once nextra ships one.
- @xmldom/xmldom ^0.9.12. speech-rule-engine 4.1.4 (via mathjax-full ->
better-react-mathjax -> nextra) pins xmldom at exactly 0.9.10, so
`npm audit fix` cannot reach it. 0.9.12 clears 13 advisories including
GHSA-6gmq-8vp8-gcm6 (high).
Docs build passes, 48 pages indexed, 0 advisories.
The published 0.7.0 bundle carries the stranded 'use client' directive
that breaks the webpack build, so the docs deploy has been failing since
the widget landed on main. 0.7.1 ships the fix.
The lockfile pinned 0.7.0 exactly and Vercel installs from the lockfile,
so publishing the fix was not enough on its own. Raising the range to
^0.7.1 keeps a fresh install off the broken build.
The widget now ships with the docs agent's embed token, so the site works
without any environment configuration. NEXT_PUBLIC_DOCSGPT_API_KEY still
overrides it for staging.
The key is public either way. NEXT_PUBLIC_* is inlined at build time, so
it lands in the client bundle whichever way it is supplied, and every
documented DocsGPT embed carries its key in the page. Abuse is bounded by
the agent's daily request and token limits, enforced by check_usage on
the answer and stream routes.
Adds a floating "Ask the docs" bubble to every docs page. docsgpt-react
was already a dependency but nothing imported it, so no widget rendered.
The widget reads window and navigator on mount, so it is lazy-loaded
with ssr: false from a client wrapper. Its light/dark theme follows the
site toggle via next-themes, promoted here from a transitive dependency
to a declared one.
The key comes from NEXT_PUBLIC_DOCSGPT_API_KEY and the host from
NEXT_PUBLIC_DOCSGPT_API_HOST. With no key set the wrapper renders
nothing rather than falling through to the demo key the package defaults
to.
Bumping docsgpt-react to ^0.7.0 also drops a duplicate React. The old
^0.6.3 pin depended on React 18 and npm installed a nested 18.3.1 beside
the app's 19.2.8, which would have thrown an invalid hook call on first
render.
pip install docsgpt now brings the web UI with it: `docsgpt api` serves the
API and the UI on one port.
- scripts/build_frontend.sh builds the frontend into docsgpt/static
(gitignored) the way the frontend image does: .env.development as the
production baseline, and index.html loading /config.js ahead of the
bundle. hatch admits the directory into the wheel and the sdist through
`artifacts`; the package workflows run the script before `uv build` and
fail if the wheel lacks the UI. The backend image keeps ignoring it.
- docsgpt/ui.py serves the build in front of Flask: files as they are,
hashed assets immutable, Flask's own path prefixes (taken from its URL map,
so new blueprints need no registration) passed through, every other GET
rendered as index.html for the client-side router. /config.js is generated
per request with VITE_API_HOST and VITE_BASE_URL set to the page's origin,
VITE_* environment variables winning. SERVE_UI=false leaves the API alone.
- docsgpt api configures gunicorn in code (gunicorn.app.base.Application)
instead of rewriting sys.argv, so the SIGUSR2 re-exec that gunicorn uses
for zero-downtime upgrades runs the docsgpt console script again and
works; verified with a live handover.
- Docs: the pip page says the UI is included, that DOCSGPT_HOME and
DOCSGPT_ENV_FILE are process environment variables rather than .env
entries, and the settings page describes SERVE_UI.
- The image pins DOCSGPT_HOME=/app: it ships no checkout, so the data home
no longer depends on the working directory.
- api, worker, beat and migrate print the data home and env file they
resolved, so an API and a worker started from different directories show
it.
- The worker passes -Q only when asked; a bare worker consumes every
configured queue, which honours EMBEDDINGS_QUEUE and DOCUMENT_PARSE_QUEUE.
- The worker runs through celery.start and returns its exit code; click
usage errors print usage and exit 2 instead of a traceback.
- Windows: solo pool and no embedded scheduler (celery rejects -B there),
with a pointer to the new `docsgpt beat` command, which runs the
scheduler on its own.
- prefetch_models and verify_offline parse their arguments, so --help is
help rather than a model name.
- A DOCSGPT_ENV_FILE that is not a file raises instead of booting with
defaults.
- pipx: replacing the CUDA torch inside the pipx environment needs
pip's --force-reinstall (pipx inject, with or without --force, leaves a
satisfied requirement alone), and --no-deps so torch's dependencies are not
reinstalled from the PyTorch index, which carries older copies of them.
- Upgrading: the embedded Milvus and LanceDB default paths now resolve under
the checkout (or DOCSGPT_HOME) instead of the start directory; the note
names the old location and the settings to pin.
- docsgpt api binds 127.0.0.1 by default, like gunicorn and uvicorn do;
--host 0.0.0.0 exposes it. The docs say so.
- The embedded Milvus and LanceDB defaults derive from the data home, so
they follow DOCSGPT_HOME like the faiss indexes and uploads do. A checkout
run from its root and the Docker image resolve to the same paths as before.
- The docs and the pyproject comment describe the CPU torch install as two
steps (torch and torchvision from the PyTorch CPU index first, then the
docling extra): pip picks the highest version across indexes, so
--extra-index-url only yields the CPU build while that index keeps pace
with PyPI.
- AGENTS.md separates DOCSGPT_HOME (moves the data home) from
DOCSGPT_ENV_FILE (selects the .env file); the docs example uses a password
placeholder.