Commit Graph
100 Commits
Author SHA1 Message Date
Alex 18313c7823 fix: keep credentials out of the text doctor borrows from its clients
Sanitising the URL in the message was not enough: the client's own error text
went into the detail too, and both psycopg and redis-py quote the URL they were
given. The reason is kept, the endpoint is kept, and the URL, username and
password are taken out of it.

The Redis test raised a generic error, so it asserted the password was absent
without ever exercising the path that leaked. Both tests now raise what the
clients actually raise.
2026-09-17 12:01:28 +01:00
Alex 9326c6edb1 fix: keep credentials out of doctor's output, and read the revision from the schema it checked
The Redis check put the whole URL in its failure message. A managed Redis URL
carries user:password@host, so a failed ping printed the password to the
terminal and into any log or issue the output was pasted into. It names
scheme://host:port/db now, and falls back to naming no URL when the value
cannot be parsed at all.

The Postgres check asked to_regclass about public.alembic_version and then read
version_num through search_path, so another schema could answer with a
different revision, or the query could fail, and doctor would send you to run
migrations against a database that is already fine.
2026-09-17 11:52:28 +01:00
Alex bcf2707efa fix: doctor reports a bad DOCSGPT_PORT instead of raising on it
_port_number was written so a hand-edited .env could not reach int() raw, and
then doctor did exactly that: a nonnumeric port ended the command with a
traceback rather than the message, in the one command whose job is to explain a
broken setup.
2026-09-17 11:43:53 +01:00
Alex 731baa7d31 test: cover what doctor actually tells you
The checks were mocked wholesale, so the branching that produces each diagnosis
had never run: a database with no schema yet, one behind this version, one that
refuses the connection, and which of the three Redis URLs failed. Each of those
is the sentence a developer reads when something is wrong, so each is pinned.

_migration_head is tested against the packaged alembic.ini itself: it needs no
database, and it is the path resolution that breaks silently when files move.
2026-09-17 11:43:06 +01:00
Alex a20c83f468 feat: a development loop in one command
`docsgpt up --native` installs services meant to outlive the shell. Development
wants the opposite, and until now it meant three terminals from the guide:
uvicorn, celery, and vite.

`docsgpt dev` runs this checkout's API and worker as children of one terminal,
both restarting when a file is saved, their output interleaved and labelled, and
Ctrl-C stopping them together. `--ui` adds the Vite dev server, `--mock-llm`
runs the bundled mock model so no API key is needed, and `--no-worker` leaves
the worker to your editor's debugger. Celery has no reloader of its own, so the
worker is wrapped in watchfiles when it is installed, and runs plain when it is
not.

Alongside it, the commands a dev loop keeps reaching for:

- `docsgpt doctor` checks what usually breaks a new setup: PostgreSQL answering
  and its schema matching this version, Redis answering, a model provider being
  configured, and the port being free.
- `docsgpt restart [api|worker]` bounces services without rewriting settings or
  rerunning migrations, which `down` plus `up` did.
- `docsgpt logs -f` follows a native install instead of telling you to run
  `tail -f` yourself.
- `docsgpt env set` applies itself to a running native install rather than
  asking you to run `docsgpt up` again to change one value.

Two bugs found on the way, both older than this change:

- `docsgpt api --reload` watched the working directory, which in a checkout is
  178,425 files: .venv, node_modules, and the indexes/ and inputs/ the app
  writes to while ingesting, so the server restarted itself mid-request. It
  watches the package now — 1,217 files.
- The VS Code "Flask Debugger" ran `flask run`, which serves only the WSGI app:
  /mcp, the SSE streams and artifact downloads 404 under it. The guide warned
  about this in prose while the debug config did it anyway. It runs uvicorn on
  the ASGI app now, like production.
2026-09-17 11:32:50 +01:00
Alex 06f233925b fix: refuse a Redis URL on port zero
parts.port returns 0 rather than raising, since 0 is inside the range it checks,
so the URL reached .env and the worker and cache had nothing to connect to.
2026-09-17 00:44:19 +01:00
Alex 51b2fa38de fix: refuse a Redis URL whose port cannot be used
urlsplit accepts an authority such as localhost:notaport or localhost:65536 and
urlunsplit rebuilds it verbatim; only parts.port raises, and nothing read it. The
unusable value reached .env, where the worker picked it up and failed to start
its broker, while `up` reported success because it waits only on the API health
endpoint.
2026-09-17 00:34:42 +01:00
Alex 147bb352ec fix: report a Redis database number too long for int() to read
The ASCII-digit check accepts any length, but since 3.11 Python refuses to
convert a digit string past its conversion limit, so a long one raised
ValueError straight through the CLI instead of the message every other
unusable URL gets.
2026-09-17 00:23:24 +01:00
Alex 0d153d3c0d fix: tie the busy-port exemption to the port the install is recorded on
The exemption asked whether any service of the install was running, so moving an
install onto a different port that something else held would pass the check and
then fail to bind, with the health poll answered by whatever owned that port —
the false success the check exists to prevent.

install.json carries the API port now, and a busy port is allowed only when it
is that port and the API service is running.
2026-09-17 00:12:40 +01:00
Alex 217310e201 fix: run the native preflights before anything is written
The Docker-stack check ran after logs/ was created, .env written and the
migrations applied, so a conflict left a migrated database and partial files
behind with no install.json — exactly the directory that down and uninstall
then refuse. It runs immediately after the port is resolved now.

Alongside it, `up --native` refuses a port it cannot bind. Neither service
manager confirms that the API bound, and /api/health carries no installation
identity, so a second install on the same port would have been answered by the
first and reported success while its own API was dead. An install re-running on
its own port is the exception, since its services are what hold it.
2026-09-17 00:03:06 +01:00
Alex 4976b3104c fix: give each native install its own services, and guard the port
From the outside-diff findings on #2800:

- Service names are derived from the install directory. A service manager has
  one namespace per user, so two installs in different --dir directories wrote
  over each other's units and down, status and uninstall acted on whichever was
  written last. The default install keeps the readable names; another directory
  gets a digest suffix.
- `up --native` refuses when a Docker stack in another directory publishes the
  same port: its API would answer the health check while these services failed
  to bind. The check degrades quietly when Docker is absent, which is exactly
  the machine a native install targets.
- A Redis database path is required to be ASCII digits: str.isdigit() is true
  for characters int() then refuses.
- DOCSGPT_PORT from a hand-edited .env is validated before conversion, and the
  error names where the bad value came from.
- Percent signs are doubled in systemd values, arguments and log paths, since
  systemd expands specifiers in all of them.
2026-09-16 23:51:20 +01:00
Alex c8ab1cbab2 fix: report a malformed Redis URL instead of raising from urlsplit
urlsplit raises ValueError on input such as redis://[::1 , which nothing
converted, so a typo left native setup with a traceback rather than the message
every other unusable URL gets.
2026-09-16 23:30:47 +01:00
Alex 65ee2f4201 fix: keep the whole Redis URL when handing out its databases
The three URLs were built by string surgery, so anything after the database
number was mangled rather than kept: rediss://host:6380/0?ssl_cert_reqs=required
came out as .../0?ssl_cert_reqs=required/0, and a URL carrying a query but no
database had /0 appended after the query. TLS and managed Redis endpoints
usually carry exactly those parameters.

The URL is split properly now, the three databases go in the path, and scheme,
credentials, host, query and fragment are preserved. A URL that cannot be
numbered this way — one that is not redis:// or rediss://, or that has
something other than a number where the database goes — is refused with a
message instead of being turned into something that merely looks like a URL.
2026-09-16 23:22:07 +01:00
Alex 76d81caaa4 fix: refuse control characters in the values of a service file
Quoting cannot carry a newline into a unit file or a plist: the line ends and
whatever follows becomes another directive. Every value bound for a service
file — the working directory, the log path, environment names and values, and
the command arguments — is checked before any of it is rendered, on both
launchd and systemd.
2026-09-16 23:10:53 +01:00
Alex 6cfd0948bc fix: refuse the Docker-only options in native mode instead of ignoring them
`up --native` took --expose, --domain and --docling and did nothing with them.
Asking for network exposure and silently getting a loopback-only install, or
asking for docling and getting an install without it, is worse than being told.
Each now says what to do instead: a reverse proxy or the Docker stack for
exposure, and the docling extra for the parser engine. --expose local and
--no-docling already describe native mode, so they stay silent.
2026-09-16 23:03:03 +01:00
Alex 501baf8aae fix: say why backup and restore do not apply to a native install
Both accepted a native install and then drove `docker compose` in a directory
with no compose file, so the user got "no configuration file provided" rather
than an explanation. They now refuse with what to do instead, and restore
refuses before it reads the archive or stops anything.
2026-09-16 23:01:30 +01:00
Alex 61f06fcce8 fix: docsgpt open points at the address a native install answers on
`open` built its address from stack.url, which honours DOCSGPT_BIND, while the
native units always bind 127.0.0.1: with a LAN bind it handed the browser an
address nothing was listening on. status had the same mismatch and was fixed
with it; both now go through one helper so they cannot drift apart again.
2026-09-16 23:00:33 +01:00
Alex 4d9f1d47a9 fix: make a native install recoverable, honest and safe to quote
From the outside-diff findings on #2800:

- install.json is written before the services are installed and started. A
  service that fails to start used to leave units behind in a directory that
  status, down and uninstall no longer recognised as a native install, so
  nothing could clean them up.
- systemd stop and removal propagate failures: `down` reporting success while
  the unit still runs, or `uninstall` dropping the unit file and the record
  while systemd still runs the service, is worse than an error. Removing a unit
  that is already gone stays harmless.
- WorkingDirectory and each Environment value are quoted and escaped for
  systemd. `--dir` takes a free-form path, and one with a space in it is not
  hypothetical: this checkout lives in one.
- Native status checks and prints http://localhost:<port>, which is what the
  units bind. With a LAN DOCSGPT_BIND it used to poll an address nothing
  listened on and call a healthy install dead.
2026-09-16 22:58:52 +01:00
Alex dbbed28f88 fix: upgrade re-execs a docsgpt it can actually find
`upgrade` exec'd the bare name `docsgpt`, so after `python -m docsgpt upgrade`
in a virtualenv without the console script on PATH, os.execv failed with a
traceback. It now uses the same launcher the service units get, which is why
that helper is no longer named for native mode.
2026-09-16 22:55:09 +01:00
Alex a96734da07 fix: run the module when the docsgpt command is not on PATH
Refusing to write the service units when `docsgpt` is not on PATH was wrong. A
package installed in a virtualenv is runnable whether or not its console script
is on PATH, and CI runs pytest as `python -m pytest`, where argv[0] is a module
file: the refusal failed thirteen native tests there.

The launcher now prefers the command on PATH, resolved to an absolute path
since PATH can hold relative entries, then an argv[0] that can be executed, and
otherwise this interpreter with `-m docsgpt`, which works wherever the package
is importable. `python -m docsgpt` became an entrypoint of its own and has a
test that runs it.
2026-09-16 22:45:54 +01:00
Alex f9f52e99f0 fix: make a native install take effect on systemd and stay out of Docker's way
From review of #2800:

- systemd `enable --now` starts nothing when the unit is already active, so a
  second `up --native` kept the old ExecStart and left the API on its previous
  port. start enables and then restarts, as the launchd path already did by
  booting the job out first.
- An explicit `home` now wins over XDG_CONFIG_HOME, which is what callers pass
  it for.
- The ExecStart program must be a real executable: when `docsgpt` is not on
  PATH, sys.argv[0] is accepted only if it can be run, and otherwise the
  failure is raised before any unit is written.
- `up --native` over a directory holding a Docker install now refuses and says
  how to proceed, instead of starting native services beside containers that
  down, status and uninstall would no longer see.
- Docs: without a terminal only --postgres-uri is required, and the Windows
  fallback names `docsgpt beat`, which the worker cannot embed there.

SystemdServices was the least covered part of the module and cannot be run on
this machine, so it now has tests for install, start, stop, remove, is_running
and a failing systemctl.
2026-09-16 22:35:14 +01:00
Alex e2c40f7a78 Merge feat/install-4-backup into feat/install-5-native 2026-09-16 22:20:46 +01:00
Alex f90c442a41 fix: cover the shutdown calls and check every volume before replacing one
Both shutdown calls sat outside the recovery that undoes them: a `compose down`
that failed partway left the stack down, and a `compose stop` that failed left
the backend and worker stopped. Each now runs inside its own try.

A restore also replaced volumes one at a time, checking each tar as it reached
it, so a damaged third payload was found with the first two already swapped in.
Every declared tar is read through first, and the imports start only once they
all come out whole.
2026-09-16 22:20:40 +01:00
Alex 9f69d39063 Merge feat/install-4-backup into feat/install-5-native 2026-09-16 22:05:26 +01:00
Alex 8cfa3fbd18 fix: let the container pick the staging directory for a restored volume
A fixed path under /tmp was both a guess about what the image can write to and
a temp-file smell that Bandit flags. The container makes the directory itself
with mktemp -d and removes it afterwards.
2026-09-16 22:05:22 +01:00
Alex 054335bf67 Merge feat/install-4-backup into feat/install-5-native 2026-09-16 22:03:52 +01:00
Alex e269743bf8 fix: start DocsGPT again when a restore fails after the stack is down
Validating the archive catches a damaged one while DocsGPT is still up, but a
well-formed archive can still hold a corrupt volume tar or a dump statement
psql refuses, and those only surface once the stack is down. The work after
the shutdown now runs inside an error boundary that starts the stack again
before the failure is reported, so a failed restore never leaves the install
stopped.
2026-09-16 22:03:48 +01:00
Alex 20f4729571 Merge feat/install-4-backup into feat/install-5-native 2026-09-16 22:01:46 +01:00
Alex 4347636496 fix: stage a restored volume under /tmp, where the image can write
The image does not run as root, so the staging directory could not be created
at the container root: `mkdir /stage` failed with permission denied and every
restore would have failed. It goes under /tmp now, and the command is built as
one string instead of concatenated pieces inside the argument list.

Checked against a real volume and the published image: a truncated tar fails
and leaves the volume exactly as it was, and a whole one restores it.
2026-09-16 22:01:39 +01:00
Alex e3c5d0511c Merge feat/install-4-backup into feat/install-5-native 2026-09-16 21:56:21 +01:00
Alex 447ae72fe2 fix: harden docsgpt backup and docsgpt restore
From review of #2799:

- The archive is created 0600 rather than at the process umask: it holds the
  install's data, and --with-settings puts .env and its secrets in it.
- restore validates everything the manifest declares before the stack is
  stopped, so a damaged archive fails while DocsGPT is still running rather
  than after `compose down` has taken it away.
- Only the volumes a backup is made of are restored. A hand-made manifest can
  no longer point import_volume at postgres_data, whose contents it empties.
- psql runs with ON_ERROR_STOP=on, so a restore that fails halfway cannot
  start DocsGPT again and call it a success.
- import_volume unpacks into the container's own filesystem first and clears
  the live volume only once the tar has come out whole, so a corrupt one
  leaves the volume as it was.
- The backend and the worker stop while the archive is made and start again
  even if the dump fails, so the dump and the volume tars describe the same
  moment instead of drifting apart as ingestion writes.
2026-09-16 21:55:33 +01:00
Alex f23a32d9c5 feat: docsgpt up --native
Run DocsGPT without Docker: the API and the worker each become a service
on the machine itself, a launchd agent on macOS and a systemd user unit
on Linux, pointed at a PostgreSQL and a Redis that already run.

`docsgpt up --native --postgres-uri ... --redis-url ...` writes the same
.env a Docker install uses, applies the migrations and starts both
services. status, logs, down and uninstall work on a native install the
same way they do on a Docker one, and never touch the database or Redis:
they were the user's to begin with.

One Redis URL covers the broker, the result backend and the cache on
three consecutive databases, starting at the one the URL names, so a
Redis that already holds something else can be shared.

Windows has neither service manager, so native mode refuses it and says
what to do instead.
2026-09-16 21:51:43 +01:00
Alex fe68fec69e feat: docsgpt backup and docsgpt restore
`docsgpt backup` writes one archive holding a pg_dump of the database, a tar
of each data volume and a manifest of what it came from; `docsgpt restore`
puts it back over an install. The settings file is left out unless
--with-settings asks for it, since it holds the install's secrets, and a
backup taken with a newer DocsGPT is refused without --force.

The volume tars go through the image the install already runs, so a backup
pulls nothing extra, and compose calls can now redirect stdout and stdin so
the dump never passes through this process.
2026-09-16 21:33:33 +01:00
Alex 4abd2c0c9f Merge pull request #2788 from arc53/feat/install-3-distribution
One-command installers: curl docs.ac/install | bash, irm docs.ac/install.ps1 | iex
2026-09-16 16:06:41 +01:00
Alex 6e452ca5cc chore: 0.21.0 2026-09-16 14:23:26 +01:00
Alex 2010af34c6 Merge pull request #2787 from arc53/feat/install-2-docsgpt-up
docsgpt up: run and manage DocsGPT on Docker from the Python package
2026-09-16 10:09:42 +01:00
Alex d2c6b5a731 Merge pull request #2786 from arc53/feat/install-1-image-compose
Serve the UI from the backend image; one-port standalone Compose stack
2026-09-16 10:09:19 +01:00
Alex d993aaced0 fix: fail on a nonzero uv installer exit; POSIX quoting for the sg handoff
The Windows installer only checked that uv.exe exists after running the uv
installer, so a failed install that left an older uv.exe behind was accepted;
it now fails on a nonzero exit code.

sg runs its command with /bin/sh, which need not be bash, so the handoff
after installing Docker quotes each argument as POSIX single quotes instead
of with bash's printf %q.
2026-09-16 01:10:11 +01:00
Alex 7e80f7a091 fix: check the uv installer against a pinned sha256 before running it
Both installers download the pinned uv installer to a file and run it only
when its sha256 matches the value pinned next to UV_VERSION; bumping the
version means bumping the hash. Astral publishes checksums for the uv
binaries but not for the installer scripts, so the hash is pinned here.

get.docker.com is still only downloaded in full before running: its content
changes over time and it publishes no checksum.
2026-09-16 01:10:11 +01:00
Alex 0e1963552c ci: each generated secret must appear exactly once 2026-09-16 01:10:11 +01:00
Alex 065adaa101 fix: the Windows installer fails when docsgpt up fails 2026-09-16 01:10:11 +01:00
Alex 28cbd268f2 ci: the installer check keeps both generated secrets 2026-09-16 01:10:11 +01:00
Alex 6afce45913 ci: the installer check requires non-empty secrets 2026-09-16 01:10:11 +01:00
Alex 49823a6859 fix: installer review follow-ups
- install.sh saves the get.docker.com and uv installers to a file and runs
  them only after the download finished, so a cut-off transfer runs nothing.
- Neither installer prints DOCSGPT_PACKAGE, which may be a URL with
  credentials.
- The CI step assigns the wheel path before exporting it, so a missing wheel
  fails instead of installing from PyPI.
- Docker-Deploying shows one code block per platform; Quickstart names the
  /opt/docsgpt home used for root on Linux.
2026-09-16 01:10:11 +01:00
Alex bd35281259 fix: plain if in the installer's uv lookup (shellcheck SC2015) 2026-09-16 01:10:11 +01:00
Alex 4f0bf2cca8 feat: one-command installers for macOS, Linux and Windows
deployment/install.sh (curl | bash) and install.ps1 (irm | iex) check for
Docker, install uv when it is missing or older than 0.8 (pinned 0.12.15 via
Astral's installer), install or upgrade the docsgpt package with
`uv tool install`, and hand the terminal to `docsgpt up` with any arguments.
On Linux without Docker the shell installer offers get.docker.com. Both run
entirely inside a function, so a download cut short runs nothing.

Releases attach both scripts next to the Compose file, which is where
docs.ac/install and docs.ac/install.ps1 will point. installer-lint.yml runs
shellcheck and the PowerShell parser; docker-image-verify.yml now installs
through install.sh. README, Quickstart, Docker-Deploying and the changelog
lead with the one-liner.
2026-09-16 01:10:11 +01:00
Alex 6b6bd1b0fb test: assert the plain-HTTP warning is printed before compose starts 2026-09-16 01:10:09 +01:00
Alex 63de66722e fix: warn about plain HTTP before the stack starts, not only afterwards
Network mode publishes the port on every interface and its access token
travels as readable text, so `docsgpt up` says so before starting rather
than in the summary at the end. The health poll's except clause says why it
swallows the error.
2026-09-16 00:53:03 +01:00
Alex db1e382502 ci: each generated secret must appear exactly once 2026-09-16 00:35:29 +01:00
Alex 9b5f1196fe ci: the repeated startup must keep both generated secrets 2026-09-16 00:20:43 +01:00
Alex 303f239fcd fix: tighten an existing .env to 0600 before rewriting it; CI checks secrets have values
envfile.update only applied 0600 when it created the file, so an existing
.env with a wider mode kept it while secrets were written into it. The mode
is now set on the open descriptor before the file is truncated and written.

The CI step now requires POSTGRES_PASSWORD and JWT_SECRET_KEY to have values:
an empty one falls back to a default without any check noticing.
2026-09-16 00:11:38 +01:00
Alex f6bf4fefd1 fix: docsgpt up review follow-ups
- wait_healthy starts no request once the deadline is reached, and neither
  its pauses nor the requests after the first run past it; the first attempt
  still always runs (status uses a zero timeout).
- envfile writes $ as $$ inside double quotes, which Compose interpolates,
  and reads $$ back as $, so a value such as pa$w'rd reaches the container
  unchanged.
- Upgrading no longer describes the working directory as the data home.
2026-09-15 23:32:37 +01:00
Alex 5e699b0168 ci: no uv cache in the image verify job (zizmor cache-poisoning) 2026-09-15 23:20:16 +01:00
Alex c7af5f873a fix: docsgpt up --adopt recreates every container
Compose keeps containers whose configuration did not change, so after a
takeover Redis and Postgres still carried the other folder's working
directory label, and the next `docsgpt up` asked for --adopt again.
2026-09-15 23:11:32 +01:00
Alex 0bd0afc1f8 fix: docsgpt up removes Caddy when leaving a domain; clearer Docker permission error
Caddy sits behind the https profile, so once COMPOSE_PROFILES no longer
enables it, `up --remove-orphans` left it running on ports 80 and 443 and
`down -v` left its volumes. `up` now removes Caddy when an install moves
off its domain, and down and uninstall name the profile explicitly.

A user outside the docker group was told Docker is not running; the error
now says how to get access to the socket.
2026-09-15 23:06:27 +01:00
Alex 68d9772f2e docs: docsgpt up and the new data home; CI runs up against the built image
Docker-Deploying gains a `docsgpt up` section, Pip-Install and Upgrading
describe the ~/.docsgpt/server data home, and the changelog covers both.

docker-image-verify.yml installs the wheel and runs `docsgpt up`, `status`,
a second `up` that must keep the secrets, and `uninstall --purge` against
the image it built. The standalone Compose file maps host.docker.internal
to the host gateway, so a model server on a Linux host is reachable the way
`docsgpt up` suggests.
2026-09-15 23:00:23 +01:00
Alex aa0e4ea280 feat: docsgpt up runs and manages DocsGPT on Docker
`docsgpt up` copies the standalone Compose file shipped with this package
version into the stack directory (~/.docsgpt/server by default), writes its
.env and starts the stack on the images of the same version. A first run
asks who should reach DocsGPT (this computer, the network with a token, or a
domain with HTTPS) and which model provider to use; flags answer the same
questions for scripts. Re-running keeps secrets and settings and moves the
image tag, and the database password is only generated for a new database.

Also: down, status, logs, token, open, env, upgrade (uv tool installs
upgrade themselves and run `up` again) and uninstall (keeps settings and
data unless --purge). The commands import no Flask, Celery or settings.

The wheel carries deployment/docker-compose-standalone.yaml as
docsgpt/deploy/docker-compose.yaml; the sdist includes the source file.
2026-09-15 22:57:00 +01:00
Alex ac06527a8d feat: installed package keeps its data in ~/.docsgpt/server
Outside a checkout the data home was the working directory, so running
`docsgpt api` from another folder silently used different settings and data.
It is now ~/.docsgpt/server (/opt/docsgpt for root on Linux); DOCSGPT_HOME
and a checkout still take precedence. The API and worker commands create
the home and point out a .env left in the working directory.
2026-09-15 22:45:54 +01:00
Alex 95fd4bdefe feat: serve the UI from the backend image, one-port standalone stack
The backend image builds the web UI with scripts/build_frontend.sh and
serves it through docsgpt/ui.py, so the standalone Compose file drops the
frontend container. UI and API share port 7091, published on 127.0.0.1
unless DOCSGPT_BIND says otherwise. POSTGRES_PASSWORD is configurable, and
an optional https profile puts Caddy in front of a public domain.

docker-image-verify.yml starts the standalone stack on the image it built
and checks the API, the UI, /config.js and a client-side route on one port.
2026-09-15 22:41:53 +01:00
Alex 0fcec68b67 Merge pull request #2784 from arc53/fix/air-gapped-review-followups
fix: air-gapped review follow-ups
2026-09-15 21:42:58 +01:00
Alex 02f2fa198a fix: disabled STT also covers live finish; clarify model cache location 2026-09-15 19:28:19 +01:00
Alex 7da46c2bea feat: air-gapped deployment guide, no implicit downloads
- Ship tiktoken's cl100k_base inside the package and build the encoding
  from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
  temp dir, and read tokenizer.json and repo metadata from that cache, so
  a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
  the endpoints return 404, audio files fail to ingest with a clear
  message, /api/config reports tts_available/stt_available, and the UI
  hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
  packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
2026-09-15 17:54:24 +01:00
Alex ca69d1ea29 Merge pull request #2780 from arc53/fix/connector-oauth-postmessage-origin
fix: stop leaking connector OAuth session tokens to other origins
2026-09-15 15:22:50 +01:00
Alex 3dfe3b6112 fix: slight hardening 2026-09-15 08:44:37 +01:00
Alex e8305ef137 fix: more mini connector hardening 2026-09-14 22:49:34 +01:00
Alex 08e8de7370 fix: mini connector fixes 2026-09-14 22:29:26 +01:00
Alex 41b3afed14 fix: stop leaking connector OAuth session tokens to other origins
The connector OAuth popup posted the session token to window.opener with
a '*' target origin, so any page that opened the popup received it. An
attacker with an account on a multi-user deployment could start a flow for
their own pending session, get a victim to finish the provider consent, and
receive a token backed by the victim's Drive/SharePoint/Confluence tokens.

- Post popup results only to allowed frontend origins: the callback origin,
  OIDC_FRONTEND_URL, the new CONNECTOR_ALLOWED_ORIGINS, and localhost:5173
  when the callback runs on a loopback host.
- Render the success page from the callback itself so the token never
  appears in a URL; callback-status ignores session_token/user_email params.
- ConnectorAuth accepts messages only from the popup it opened, on the
  callback origin reported by /api/connectors/auth.
- /api/connectors/disconnect requires auth and only deletes the caller's
  session.
- /api/connectors/sync and /api/remote reject session tokens the caller
  does not own.

Fixes #2766
2026-09-14 21:46:09 +01:00
Alex d121eec2f5 Merge pull request #2771 from arc53/fix/durable-non-text-uploads
fix: handle non text uploads more carefully
2026-09-14 21:23:22 +01:00
Alex ad201f8318 fix: refuse oversized TIFF/BMP attachments before converting them
A deflate-compressed TIFF under 1 MB can declare 144 million pixels and
take 1.2 GB to convert to PNG, and Pillow only warns below 179 million.
Read the dimensions from the header and refuse images over 40 million
pixels before any pixel data is decoded. Pillow's DecompressionBombError
is now raised as DocumentParseError, so the upload fails once instead of
being retried.
2026-09-14 17:45:43 +01:00
Alex 3030aba39d fix: handle non text uploads more carefully 2026-09-14 17:27:13 +01:00
Alex 82691314ca Merge pull request #2769 from arc53/feat/async-starlette-streams
feat(api): serve long-lived streams on the event loop
2026-09-14 16:22:36 +01:00
Alex 1bf7380aa5 fix: mini test 2026-09-14 14:12:21 +01:00
Alex 72904b0b95 fix: Make ticket claim atomic. 2026-09-14 14:06:59 +01:00
Alex fb66a0b230 fix: stability improvements on event loop 2026-09-14 12:52:15 +01:00
Alex e889392eba fix(api): address review on the event-loop streams
- Per-user SSE cap uses per-connection leases in a sorted set instead of
  a shared INCR/DECR counter with a TTL. A stream that outlived the TTL
  could let the counter expire, then decrement another stream's slot or
  drive the count negative. Leases refresh while a stream sends frames
  and age out when a stream dies without cleanup.
- Shielded cleanup is bounded per step: stream close and on_close in
  ClosingStreamingResponse, unsubscribe and close in AsyncTopic, lease
  release, and the replay-budget check, so a dead Redis connection can't
  hold a request or a graceful shutdown.
- Artifact downloads send an ASCII filename plus an RFC 5987 filename*
  for non-ASCII names, build headers before opening the file, and close
  disk handles off the event loop.
- Type hints and docstrings on the new helpers.
2026-09-14 10:19:59 +01:00
Alex 54b540ea42 feat(api): serve long-lived streams on the event loop
Move GET /api/events, the remote-device command stream and artifact
downloads from Flask to Starlette routes mounted ahead of the Flask
catch-all. On Flask each held an a2wsgi threadpool slot for as long as
its response stayed open, and because uvicorn drops writes after a
client disconnects, a closed tab never released it.

- asgi_auth: one JWT/OIDC gate for Starlette routes; the chat reconnect
  reader uses it too
- ClosingStreamingResponse closes the body iterator and releases the
  SSE slot or file handle even when the client leaves before the first
  frame
- AsyncTopic liveness probe replaces the sync client's socket_timeout
  guard against half-open pub/sub sockets
- ASYNC_REDIS_MAX_CONNECTIONS sizes the async Redis pool; every open
  stream holds a connection and redis-py defaults to 100
2026-09-14 08:33:46 +01:00
Alex 38773471a8 Update docsgpt.yaml 2026-09-13 14:48:28 +01:00
Alex eae635b6dd Merge pull request #2765 from arc53/fix/release-recovery-and-attestations
ci: fix PyPI publishing from the release chain, make image builds recoverable
2026-09-12 23:50:26 +01:00
Alex 12016baaf2 ci: resolve the manual version under refs/tags
The version input reached checkout as a bare ref, so a manual run given a
branch name would have built that branch and published it as an image tag,
`latest` included. It also meant a version that happens to match a branch name
would silently resolve to the branch rather than the tag.

Resolve it under refs/tags/ instead. Plain versions from backend-release keep
working, and anything that is not a tag fails at checkout.
2026-09-12 23:41:43 +01:00
Alex 99b8508bdc ci: build the source the version names
A manual run takes the version as an input but the checkout had no ref, so it
built whatever the run was dispatched from. Dispatching off main to rebuild an
older release would have published current main under that release's tags, and
moved `latest` to it.

Check out the version instead, falling back to the run's own ref when there is
no version, which is the push that publishes `develop`. The release and call
paths already checked out the right commit; this makes them explicit and fails
loudly on a version with no tag rather than mislabelling an image. The compose
file attached to the release now comes from the tag too.
2026-09-12 23:37:14 +01:00
Alex 55e619b812 ci: let the image workflows be run by hand
None of the three image workflows could be started manually, so when 0.20.0
tagged before the release chain knew about the frontend and sandbox images,
there was no way to build those images for the tag short of recreating the
release. Add a workflow_dispatch trigger taking the version to build.

A manual run also gets a `move_latest` switch, so rebuilding an older tag does
not drag `latest` backwards. The rule for the moving tag now lives in the
manifest step's shell rather than in a nested expression: `latest` follows the
build unless the release is a prerelease, the run asked it not to, or it is the
rolling `develop` build.
2026-09-12 23:29:33 +01:00
Alex 15c4d82bcf ci: stop attesting when the release workflow publishes to PyPI
0.20.0 failed to upload with "Certificate's Build Config URI ... does not
match expected Trusted Publisher". Authentication was fine: trusted publishing
matches the called workflow, pypi-publish.yml, which is what PyPI is
configured with. The PEP 740 attestation the publish action attaches by
default records the workflow that STARTED the run instead, backend-release.yml,
and PyPI verifies that against the same publisher entry. The two claims can
never agree while the publish runs as a called workflow, and no publisher
configuration satisfies both.

Attest only when this file is the entry point, which covers the manual
dispatch and a release published by hand. The release chain uploads unattested
rather than failing.

The earlier rehearsal went to TestPyPI and passed, so this only appeared on the
real index.
2026-09-12 23:29:33 +01:00
Alex c60b87ac12 Merge pull request #2763 from arc53/fix/release-frontend-image
release: publish the frontend and sandbox images, write the 0.20.0 changelog
2026-09-12 23:02:08 +01:00
Alex c36b0af170 chore: bump version to 0.20.0 2026-09-12 22:28:23 +01:00
Alex c18e26f798 ci: keep prereleases off the latest tag
`release: published` fires for prereleases too, and every manifest job pushed
`latest` unconditionally, so publishing a release candidate by hand would have
made it the image everyone pulls. The backend image had the same hole, so fix
all three rather than only the two added here.

A release the backend-release workflow calls is always stable, so the
workflow_call path keeps moving `latest`; the release-event path moves it only
when the release is not a prerelease.
2026-09-12 22:19:27 +01:00
Alex 7dd59519be docs: write the 0.20.0 changelog
The changelog page has been an empty stub carrying only a title, and was
hidden from the sidebar. Fill it in from the 38 pull requests merged since the
0.19.0 tag, grouped by what a reader would look for, and link the GitHub
release notes for the full per-PR list and the upgrade guide for the steps an
existing deployment has to take. Earlier releases are not backfilled.

Now that the page has content, show it in the sidebar.
2026-09-12 22:06:28 +01:00
Alex fc5992c5d4 ci: publish the docsgpt-sandbox image
deployment/k8s/deployments/sandbox-deploy.yaml pulls arc53/docsgpt-sandbox,
which has never been pushed anywhere: Compose builds the runner from the
checkout (`build: ./sandbox`), but Kubernetes cannot build, so enabling code
execution on a cluster failed on an image that does not exist.

Build and push it like the other two images: `develop` on a push to main that
touches deployment/sandbox, and `<version>` plus `latest` when the release
workflow calls it. Release and develop live in one file here rather than two,
because the runner changes rarely and the only difference is which tags move.
The tag comes from the inputs and the release payload, not from
`github.event_name`, which is `push` when backend-release calls this.
2026-09-12 22:06:27 +01:00
Alex 01dfe473d3 ci: publish the frontend image from the release chain
backend-release.yml creates the GitHub release with GITHUB_TOKEN, and GitHub
never starts workflows from events that token produces, so the frontend image
workflow's `release: published` trigger does not fire for a version bump on
main. It was the only publish workflow left out of the workflow_call fix, so
the next bump would push arc53/docsgpt:<version> and move docsgpt:latest while
docsgpt-fe stayed on the previous release — and the standalone compose file
defaults both images to latest.

Give cife.yml a workflow_call trigger with a version input (same shape as
ci.yml) and call it from backend-release.yml after the release is created. The
release trigger stays for releases created by hand.

While here, bring it in line with ci.yml: run the publishing jobs in the
docker-hub environment, tag from the public arc53 namespace instead of the
login secret, pin the actions by digest, add the OCI version label, and drop
the QEMU step that only ran on the native arm64 runner.
2026-09-12 20:28:58 +01:00
Alex b2a5480dab Merge pull request #2762 from arc53/chore/bump-frontend-deps-2026-09
chore(deps): bump frontend, docs, widget and e2e dependencies
2026-09-12 20:13:04 +01:00
Alex 4cc603db89 fix(frontend): guard CopyButton against concurrent copies
The button is disabled once isCopied is true, but copy() is now async, so
isCopied and the disabled prop only catch up after it resolves. Clicks
landing inside that window both passed the guard and started a write.

Nothing leaked, since the existing clearTimeout already handled the
duplicate timer, but the window did not exist before copy-to-clipboard 4
made the call async. A ref cleared in finally closes it.

Raised by CodeRabbit on #2762. 598 tests pass, build and lint unchanged.
2026-09-12 20:05:20 +01:00
Alex 88b4d89b4f chore(deps): e2e TypeScript 7
TypeScript 7 removed the baseUrl compiler option, so the path mapping is
rewritten relative ("./helpers/*"), which is what the removal notice
prescribes.

tsc --noEmit is clean and `playwright test --list` collects all 221 tests
across 42 files, so module resolution still works at runtime.
2026-09-12 19:35:22 +01:00
Alex 05c53195a3 chore(deps): react-widget majors -- markdown-it 15, babel 8, flow-bin 0.331
markdown-it 15 renders byte-identical HTML to 14.3.2 across headings,
lists, code, links, tables, blockquotes, raw HTML, javascript: URLs,
entities and rules, with the widget's custom table/thead/tr/td/th renderer
rules applied. Output still goes through DOMPurify before injection.

Babel 8 and flow-bin change nothing in the build: there is no babel config
in the package and Parcel transforms with SWC, so bundle sizes are
identical before and after. They are bumped to stay current.

Verified the built bundle in Chromium: modern/browser.js loads with no page
errors, renderDocsGPTWidget and renderSearchBar are both functions, and
both mount DOM.

ESLint 10 is not taken here either -- same eslint-plugin-react block as the
frontend, so eslint and @eslint/js stay on 9.

Build, tsc --noEmit and lint all clean, 0 advisories.
2026-09-12 19:34:51 +01:00
Alex def32a44fc chore(deps): declare eslint config deps, eslint-plugin-n 18
eslint.config.js imports @eslint/js and globals but neither was declared;
both resolved transitively through eslint. Declaring them fixes a latent
break and is a prerequisite for ESLint 10, which stops providing them.

ESLint 10 itself is not taken. eslint-plugin-react 7.37.5 is the latest
release and peers at eslint "^3 || ... || ^9.7", and under 10 every rule
load throws "contextOrFilename.getFilename is not a function" because
ESLint 10 removed context.getFilename(). The frontend stays on eslint 9
until that ships; note npm already marks 9.x as no longer supported.

eslint-plugin-n goes to 18, though the flat config does not reference it
(nor eslint-plugin-import, eslint-plugin-promise or eslint-config-prettier
-- all four are unused devDependencies).

Lint output unchanged from baseline.
2026-09-12 19:32:49 +01:00
Alex 05d9338983 chore(deps): mermaid 12
mermaid 12 needs no source change; mermaidSecurity.ts and its strict-mode
init-directive test pass unchanged.

Adds a lodash-es ^4.18.1 override. mermaid 12 depends on chevrotain ~11.1.2
(mermaid 11 resolved chevrotain 12, which has no lodash-es), and chevrotain
11 pulls lodash-es <= 4.17.23, carrying GHSA-r5fr-rjxr-66jc and
GHSA-f23m-r3pf-42rh, both high. 4.18.1 is the fixed line. Without the
override the frontend gains 5 high-severity advisories.

Verified rendering in Chromium rather than the unit suite, since happy-dom
has no SVG getBBox and cannot render any mermaid version. Flowchart,
sequence, pie, class, state, gantt, ER and mindmap all produce SVG on 12,
matching 11 to the byte except classDiagram (+17 bytes). A strict-mode
`click ... "javascript:..."` target is still stripped.

598 tests pass, build clean, 0 advisories.
2026-09-12 19:31:29 +01:00
Alex 6319b25b56 chore(deps): react-markdown 10
react-markdown 10 removed the `className` prop and no longer renders a
wrapper element; passing className now throws at render time. Each of the
three call sites gets an explicit wrapping div carrying the same classes,
which reproduces the v9 DOM exactly.

The large diff is re-indentation from the added wrapper. `git diff -w`
shows the only substantive change per file is the moved className.

Fixes the 19 tests that failed on the upgrade. 598 tests pass, build clean,
lint unchanged.
2026-09-12 19:28:18 +01:00
Alex cf1c9d542d chore(deps): copy-to-clipboard 4, vite-plugin-svgr 5, @shadcn/react 0.3.1
copy-to-clipboard 4 returns Promise<boolean> instead of boolean, so
CopyButton's handleCopy awaits it. Without the await the success branch
always ran, since a Promise is always truthy. Awaiting also routes a
rejected copy into the existing catch.

vite-plugin-svgr 5 and @shadcn/react 0.3.1 needed no source changes.

598 tests pass, build clean, lint unchanged.
2026-09-12 19:26:20 +01:00
Alex 08a2c038bc chore(deps): react-widget in-range updates
Babel 7.29.7, React 19.3, styled-components 6.5.3, dompurify 3.4.15,
typescript-eslint 8.70, prettier 3.9.6, svgo 4.1, markdown-it 14.3.

Clears all 5 npm audit advisories, including svgo removeScripts
GHSA-w27v-7q3p-w38r and GHSA-4vpr-x523-8j87 (both high).

Prettier 3.9 changed ternary and union-type wrapping, so two src files are
reformatted by lint-fix. No behavior change.

Build, tsc --noEmit and lint all clean.
2026-09-12 19:25:00 +01:00
Alex d1f75b1d82 chore(deps): e2e in-range updates
Playwright 1.63, pg 8.23, @types/node 22.20. tsc --noEmit clean.
2026-09-12 19:23:07 +01:00
Alex ca889de288 chore(deps): docs in-range updates, pin zod, override xmldom
Next 16.3.5 and React 19.3 in docs/, plus the transitive refresh that came
with them (mermaid 11.17, katex 0.16.47, shiki 3.23, styled-components 6.5).

Two overrides were needed:

- zod ~4.3.6. nextra 4.6.1 declares `zod: ^4.1.12` but its MDX prop
  validation breaks on zod >= 4.4.0, failing prerender for every page with
  "Invalid input: expected nonoptional, received undefined -> at children".
  Bisected: 4.3.6 builds, 4.4.0 through 4.6.2 all fail. nextra 4.6.1 is the
  latest release, so there is no upstream fix to take yet. Drop this pin
  once nextra ships one.

- @xmldom/xmldom ^0.9.12. speech-rule-engine 4.1.4 (via mathjax-full ->
  better-react-mathjax -> nextra) pins xmldom at exactly 0.9.10, so
  `npm audit fix` cannot reach it. 0.9.12 clears 13 advisories including
  GHSA-6gmq-8vp8-gcm6 (high).

Docs build passes, 48 pages indexed, 0 advisories.
2026-09-12 19:23:07 +01:00
Alex 19ad1cb828 chore(deps): frontend in-range updates
Applies all semver-compatible updates in frontend/: React 19.3, Vite 8.3,
vitest 4.1.11, radix-ui 1.6.7, lucide-react 1.45, mermaid 11.17,
typescript-eslint 8.70, tailwind 4.3.3, i18next 26.4 and others.

Clears all 4 npm audit advisories (js-yaml GHSA-2883-xcg3-v3hh high,
@vitest/mocker GHSA-82fw-gwwq-j7x9, @humanfs/node GHSA-p498-v437-472g).

598 tests pass, build clean, lint unchanged from baseline.
2026-09-12 19:16:35 +01:00