networkx.pagerank delegates to a scipy implementation, and scipy is not a
DocsGPT dependency — it only arrives transitively through the optional
docling extra. In a default install every graph retrieval raised
ModuleNotFoundError inside _ppr_scores, hit the per-source except, and
degraded to ClassicRAG: the graph was built and paid for, then never used,
with one ERROR line per source per query as the only signal.
Rank with a local power iteration over the same row-normalized transition
matrix: undirected edges normalized per endpoint, dangling nodes
redistributed along the restart vector, and the restart vector normalized
across the nodes the subgraph actually holds so seed mass cannot leak.
Parity with networkx is asserted while scipy happens to be installed in the
test env, and the retrieval path is exercised with the import blocked.
_redis_urls raises before any check runs and cli.main prints what it raises, so
every one of its five messages interpolated the URL it was given — password and
all — straight to stderr. They render it through _endpoint now.
_endpoint itself then echoed the whole path, so an over-long database number
came back at full length: 5132 characters of error for one bad setting. It caps
the path it renders, which protects every caller, since that string is written
into terminals, CI logs and error messages rather than reused as a URL.
CodeQL flags "host" in message as incomplete URL sanitisation. I removed that
shape from two assertions and introduced a third in the same commit; this takes
it out of the provider test and the remaining Redis one, comparing with what
_endpoint produced instead, which is the contract those tests actually mean.
- doctor printed OPENAI_BASE_URL raw, the third place a credential-bearing URL
reached the terminal; it goes through _endpoint like the others.
- dev.run left its output readers unjoined, so a child's last lines could be
lost on exit. The threads are kept and joined during teardown.
- logs read each file and then reopened it to follow, so anything written in
between appeared in neither. One handle now serves both.
- dev took --port straight from argparse: 0 would have served on an ephemeral
port while printing 0, and oversized values reach socket.bind. It goes
through _port_number first.
- doctor --redis-url overrode only the broker, so a stale result backend or
cache was still pinged and the flag looked broken. All three endpoints now
come from the URL given.
CodeQL flags `"host:port" in message` as incomplete URL sanitisation. It is a
message rather than a URL being authorised, so it is not a vulnerability, but
the check was red and comparing against the sanitiser's own output asserts the
real contract. The signal handler now says why it swallows: the child is
already gone, and shutdown must not fail on what it is cleaning up.
env.py sets no version_table_schema, so the table follows search_path. Pinning
public. in doctor's queries made them agree with each other and disagree with
alembic: on an install using another schema it reported no schema at all and
sent the user to migrate an already-migrated database. Both queries resolve the
table the same way alembic does now.
Both preflights passed for `dev --mock-llm --port 8090`: the port was free, and
then the mock and the API were each handed it. One child could not bind, and the
API was pointed at that port as its model server while trying to listen on it.
The mock starts first and the API and worker are pointed at it, but only the
API port was preflighted. A busy 8090 therefore showed up as a child exiting
once the rest were running, or as the API talking to whatever else was on that
port. --ui is deliberately left alone: vite.config.ts sets no strictPort, so
Vite moves to the next free port rather than failing.
urlsplit accepts redis://host:notaport/0; only parts.port raises, and it raises
on access rather than at split time, so the ValueError fell outside the try.
_check_redis catches the client error and then formats it through _endpoint, so
doctor ended with a traceback from inside its own error path.
Running outside a checkout and starting on a port something else holds are both
refusals a developer will meet, and neither had a test. The second also asserts
that nothing is spawned when the port is taken, which is the part that matters:
the guard runs before any child process exists.
Sanitising the URL in the message was not enough: the client's own error text
went into the detail too, and both psycopg and redis-py quote the URL they were
given. The reason is kept, the endpoint is kept, and the URL, username and
password are taken out of it.
The Redis test raised a generic error, so it asserted the password was absent
without ever exercising the path that leaked. Both tests now raise what the
clients actually raise.
The Redis check put the whole URL in its failure message. A managed Redis URL
carries user:password@host, so a failed ping printed the password to the
terminal and into any log or issue the output was pasted into. It names
scheme://host:port/db now, and falls back to naming no URL when the value
cannot be parsed at all.
The Postgres check asked to_regclass about public.alembic_version and then read
version_num through search_path, so another schema could answer with a
different revision, or the query could fail, and doctor would send you to run
migrations against a database that is already fine.
_port_number was written so a hand-edited .env could not reach int() raw, and
then doctor did exactly that: a nonnumeric port ended the command with a
traceback rather than the message, in the one command whose job is to explain a
broken setup.
The checks were mocked wholesale, so the branching that produces each diagnosis
had never run: a database with no schema yet, one behind this version, one that
refuses the connection, and which of the three Redis URLs failed. Each of those
is the sentence a developer reads when something is wrong, so each is pinned.
_migration_head is tested against the packaged alembic.ini itself: it needs no
database, and it is the path resolution that breaks silently when files move.
`docsgpt up --native` installs services meant to outlive the shell. Development
wants the opposite, and until now it meant three terminals from the guide:
uvicorn, celery, and vite.
`docsgpt dev` runs this checkout's API and worker as children of one terminal,
both restarting when a file is saved, their output interleaved and labelled, and
Ctrl-C stopping them together. `--ui` adds the Vite dev server, `--mock-llm`
runs the bundled mock model so no API key is needed, and `--no-worker` leaves
the worker to your editor's debugger. Celery has no reloader of its own, so the
worker is wrapped in watchfiles when it is installed, and runs plain when it is
not.
Alongside it, the commands a dev loop keeps reaching for:
- `docsgpt doctor` checks what usually breaks a new setup: PostgreSQL answering
and its schema matching this version, Redis answering, a model provider being
configured, and the port being free.
- `docsgpt restart [api|worker]` bounces services without rewriting settings or
rerunning migrations, which `down` plus `up` did.
- `docsgpt logs -f` follows a native install instead of telling you to run
`tail -f` yourself.
- `docsgpt env set` applies itself to a running native install rather than
asking you to run `docsgpt up` again to change one value.
Two bugs found on the way, both older than this change:
- `docsgpt api --reload` watched the working directory, which in a checkout is
178,425 files: .venv, node_modules, and the indexes/ and inputs/ the app
writes to while ingesting, so the server restarted itself mid-request. It
watches the package now — 1,217 files.
- The VS Code "Flask Debugger" ran `flask run`, which serves only the WSGI app:
/mcp, the SSE streams and artifact downloads 404 under it. The guide warned
about this in prose while the debug config did it anyway. It runs uvicorn on
the ASGI app now, like production.
parts.port returns 0 rather than raising, since 0 is inside the range it checks,
so the URL reached .env and the worker and cache had nothing to connect to.
urlsplit accepts an authority such as localhost:notaport or localhost:65536 and
urlunsplit rebuilds it verbatim; only parts.port raises, and nothing read it. The
unusable value reached .env, where the worker picked it up and failed to start
its broker, while `up` reported success because it waits only on the API health
endpoint.
The ASCII-digit check accepts any length, but since 3.11 Python refuses to
convert a digit string past its conversion limit, so a long one raised
ValueError straight through the CLI instead of the message every other
unusable URL gets.
The exemption asked whether any service of the install was running, so moving an
install onto a different port that something else held would pass the check and
then fail to bind, with the health poll answered by whatever owned that port —
the false success the check exists to prevent.
install.json carries the API port now, and a busy port is allowed only when it
is that port and the API service is running.
The Docker-stack check ran after logs/ was created, .env written and the
migrations applied, so a conflict left a migrated database and partial files
behind with no install.json — exactly the directory that down and uninstall
then refuse. It runs immediately after the port is resolved now.
Alongside it, `up --native` refuses a port it cannot bind. Neither service
manager confirms that the API bound, and /api/health carries no installation
identity, so a second install on the same port would have been answered by the
first and reported success while its own API was dead. An install re-running on
its own port is the exception, since its services are what hold it.
From the outside-diff findings on #2800:
- Service names are derived from the install directory. A service manager has
one namespace per user, so two installs in different --dir directories wrote
over each other's units and down, status and uninstall acted on whichever was
written last. The default install keeps the readable names; another directory
gets a digest suffix.
- `up --native` refuses when a Docker stack in another directory publishes the
same port: its API would answer the health check while these services failed
to bind. The check degrades quietly when Docker is absent, which is exactly
the machine a native install targets.
- A Redis database path is required to be ASCII digits: str.isdigit() is true
for characters int() then refuses.
- DOCSGPT_PORT from a hand-edited .env is validated before conversion, and the
error names where the bad value came from.
- Percent signs are doubled in systemd values, arguments and log paths, since
systemd expands specifiers in all of them.
urlsplit raises ValueError on input such as redis://[::1 , which nothing
converted, so a typo left native setup with a traceback rather than the message
every other unusable URL gets.
The three URLs were built by string surgery, so anything after the database
number was mangled rather than kept: rediss://host:6380/0?ssl_cert_reqs=required
came out as .../0?ssl_cert_reqs=required/0, and a URL carrying a query but no
database had /0 appended after the query. TLS and managed Redis endpoints
usually carry exactly those parameters.
The URL is split properly now, the three databases go in the path, and scheme,
credentials, host, query and fragment are preserved. A URL that cannot be
numbered this way — one that is not redis:// or rediss://, or that has
something other than a number where the database goes — is refused with a
message instead of being turned into something that merely looks like a URL.
Quoting cannot carry a newline into a unit file or a plist: the line ends and
whatever follows becomes another directive. Every value bound for a service
file — the working directory, the log path, environment names and values, and
the command arguments — is checked before any of it is rendered, on both
launchd and systemd.
`up --native` took --expose, --domain and --docling and did nothing with them.
Asking for network exposure and silently getting a loopback-only install, or
asking for docling and getting an install without it, is worse than being told.
Each now says what to do instead: a reverse proxy or the Docker stack for
exposure, and the docling extra for the parser engine. --expose local and
--no-docling already describe native mode, so they stay silent.
Both accepted a native install and then drove `docker compose` in a directory
with no compose file, so the user got "no configuration file provided" rather
than an explanation. They now refuse with what to do instead, and restore
refuses before it reads the archive or stops anything.
`open` built its address from stack.url, which honours DOCSGPT_BIND, while the
native units always bind 127.0.0.1: with a LAN bind it handed the browser an
address nothing was listening on. status had the same mismatch and was fixed
with it; both now go through one helper so they cannot drift apart again.
From the outside-diff findings on #2800:
- install.json is written before the services are installed and started. A
service that fails to start used to leave units behind in a directory that
status, down and uninstall no longer recognised as a native install, so
nothing could clean them up.
- systemd stop and removal propagate failures: `down` reporting success while
the unit still runs, or `uninstall` dropping the unit file and the record
while systemd still runs the service, is worse than an error. Removing a unit
that is already gone stays harmless.
- WorkingDirectory and each Environment value are quoted and escaped for
systemd. `--dir` takes a free-form path, and one with a space in it is not
hypothetical: this checkout lives in one.
- Native status checks and prints http://localhost:<port>, which is what the
units bind. With a LAN DOCSGPT_BIND it used to poll an address nothing
listened on and call a healthy install dead.
`upgrade` exec'd the bare name `docsgpt`, so after `python -m docsgpt upgrade`
in a virtualenv without the console script on PATH, os.execv failed with a
traceback. It now uses the same launcher the service units get, which is why
that helper is no longer named for native mode.
Refusing to write the service units when `docsgpt` is not on PATH was wrong. A
package installed in a virtualenv is runnable whether or not its console script
is on PATH, and CI runs pytest as `python -m pytest`, where argv[0] is a module
file: the refusal failed thirteen native tests there.
The launcher now prefers the command on PATH, resolved to an absolute path
since PATH can hold relative entries, then an argv[0] that can be executed, and
otherwise this interpreter with `-m docsgpt`, which works wherever the package
is importable. `python -m docsgpt` became an entrypoint of its own and has a
test that runs it.
From review of #2800:
- systemd `enable --now` starts nothing when the unit is already active, so a
second `up --native` kept the old ExecStart and left the API on its previous
port. start enables and then restarts, as the launchd path already did by
booting the job out first.
- An explicit `home` now wins over XDG_CONFIG_HOME, which is what callers pass
it for.
- The ExecStart program must be a real executable: when `docsgpt` is not on
PATH, sys.argv[0] is accepted only if it can be run, and otherwise the
failure is raised before any unit is written.
- `up --native` over a directory holding a Docker install now refuses and says
how to proceed, instead of starting native services beside containers that
down, status and uninstall would no longer see.
- Docs: without a terminal only --postgres-uri is required, and the Windows
fallback names `docsgpt beat`, which the worker cannot embed there.
SystemdServices was the least covered part of the module and cannot be run on
this machine, so it now has tests for install, start, stop, remove, is_running
and a failing systemctl.
Both shutdown calls sat outside the recovery that undoes them: a `compose down`
that failed partway left the stack down, and a `compose stop` that failed left
the backend and worker stopped. Each now runs inside its own try.
A restore also replaced volumes one at a time, checking each tar as it reached
it, so a damaged third payload was found with the first two already swapped in.
Every declared tar is read through first, and the imports start only once they
all come out whole.
A fixed path under /tmp was both a guess about what the image can write to and
a temp-file smell that Bandit flags. The container makes the directory itself
with mktemp -d and removes it afterwards.
Validating the archive catches a damaged one while DocsGPT is still up, but a
well-formed archive can still hold a corrupt volume tar or a dump statement
psql refuses, and those only surface once the stack is down. The work after
the shutdown now runs inside an error boundary that starts the stack again
before the failure is reported, so a failed restore never leaves the install
stopped.
The image does not run as root, so the staging directory could not be created
at the container root: `mkdir /stage` failed with permission denied and every
restore would have failed. It goes under /tmp now, and the command is built as
one string instead of concatenated pieces inside the argument list.
Checked against a real volume and the published image: a truncated tar fails
and leaves the volume exactly as it was, and a whole one restores it.
From review of #2799:
- The archive is created 0600 rather than at the process umask: it holds the
install's data, and --with-settings puts .env and its secrets in it.
- restore validates everything the manifest declares before the stack is
stopped, so a damaged archive fails while DocsGPT is still running rather
than after `compose down` has taken it away.
- Only the volumes a backup is made of are restored. A hand-made manifest can
no longer point import_volume at postgres_data, whose contents it empties.
- psql runs with ON_ERROR_STOP=on, so a restore that fails halfway cannot
start DocsGPT again and call it a success.
- import_volume unpacks into the container's own filesystem first and clears
the live volume only once the tar has come out whole, so a corrupt one
leaves the volume as it was.
- The backend and the worker stop while the archive is made and start again
even if the dump fails, so the dump and the volume tars describe the same
moment instead of drifting apart as ingestion writes.
Run DocsGPT without Docker: the API and the worker each become a service
on the machine itself, a launchd agent on macOS and a systemd user unit
on Linux, pointed at a PostgreSQL and a Redis that already run.
`docsgpt up --native --postgres-uri ... --redis-url ...` writes the same
.env a Docker install uses, applies the migrations and starts both
services. status, logs, down and uninstall work on a native install the
same way they do on a Docker one, and never touch the database or Redis:
they were the user's to begin with.
One Redis URL covers the broker, the result backend and the cache on
three consecutive databases, starting at the one the URL names, so a
Redis that already holds something else can be shared.
Windows has neither service manager, so native mode refuses it and says
what to do instead.
`docsgpt backup` writes one archive holding a pg_dump of the database, a tar
of each data volume and a manifest of what it came from; `docsgpt restore`
puts it back over an install. The settings file is left out unless
--with-settings asks for it, since it holds the install's secrets, and a
backup taken with a newer DocsGPT is refused without --force.
The volume tars go through the image the install already runs, so a backup
pulls nothing extra, and compose calls can now redirect stdout and stdin so
the dump never passes through this process.
The Windows installer only checked that uv.exe exists after running the uv
installer, so a failed install that left an older uv.exe behind was accepted;
it now fails on a nonzero exit code.
sg runs its command with /bin/sh, which need not be bash, so the handoff
after installing Docker quotes each argument as POSIX single quotes instead
of with bash's printf %q.
Both installers download the pinned uv installer to a file and run it only
when its sha256 matches the value pinned next to UV_VERSION; bumping the
version means bumping the hash. Astral publishes checksums for the uv
binaries but not for the installer scripts, so the hash is pinned here.
get.docker.com is still only downloaded in full before running: its content
changes over time and it publishes no checksum.
- install.sh saves the get.docker.com and uv installers to a file and runs
them only after the download finished, so a cut-off transfer runs nothing.
- Neither installer prints DOCSGPT_PACKAGE, which may be a URL with
credentials.
- The CI step assigns the wheel path before exporting it, so a missing wheel
fails instead of installing from PyPI.
- Docker-Deploying shows one code block per platform; Quickstart names the
/opt/docsgpt home used for root on Linux.
deployment/install.sh (curl | bash) and install.ps1 (irm | iex) check for
Docker, install uv when it is missing or older than 0.8 (pinned 0.12.15 via
Astral's installer), install or upgrade the docsgpt package with
`uv tool install`, and hand the terminal to `docsgpt up` with any arguments.
On Linux without Docker the shell installer offers get.docker.com. Both run
entirely inside a function, so a download cut short runs nothing.
Releases attach both scripts next to the Compose file, which is where
docs.ac/install and docs.ac/install.ps1 will point. installer-lint.yml runs
shellcheck and the PowerShell parser; docker-image-verify.yml now installs
through install.sh. README, Quickstart, Docker-Deploying and the changelog
lead with the one-liner.
Network mode publishes the port on every interface and its access token
travels as readable text, so `docsgpt up` says so before starting rather
than in the summary at the end. The health poll's except clause says why it
swallows the error.
envfile.update only applied 0600 when it created the file, so an existing
.env with a wider mode kept it while secrets were written into it. The mode
is now set on the open descriptor before the file is truncated and written.
The CI step now requires POSTGRES_PASSWORD and JWT_SECRET_KEY to have values:
an empty one falls back to a default without any check noticing.
- wait_healthy starts no request once the deadline is reached, and neither
its pauses nor the requests after the first run past it; the first attempt
still always runs (status uses a zero timeout).
- envfile writes $ as $$ inside double quotes, which Compose interpolates,
and reads $$ back as $, so a value such as pa$w'rd reaches the container
unchanged.
- Upgrading no longer describes the working directory as the data home.
Compose keeps containers whose configuration did not change, so after a
takeover Redis and Postgres still carried the other folder's working
directory label, and the next `docsgpt up` asked for --adopt again.
Caddy sits behind the https profile, so once COMPOSE_PROFILES no longer
enables it, `up --remove-orphans` left it running on ports 80 and 443 and
`down -v` left its volumes. `up` now removes Caddy when an install moves
off its domain, and down and uninstall name the profile explicitly.
A user outside the docker group was told Docker is not running; the error
now says how to get access to the socket.
Docker-Deploying gains a `docsgpt up` section, Pip-Install and Upgrading
describe the ~/.docsgpt/server data home, and the changelog covers both.
docker-image-verify.yml installs the wheel and runs `docsgpt up`, `status`,
a second `up` that must keep the secrets, and `uninstall --purge` against
the image it built. The standalone Compose file maps host.docker.internal
to the host gateway, so a model server on a Linux host is reachable the way
`docsgpt up` suggests.
`docsgpt up` copies the standalone Compose file shipped with this package
version into the stack directory (~/.docsgpt/server by default), writes its
.env and starts the stack on the images of the same version. A first run
asks who should reach DocsGPT (this computer, the network with a token, or a
domain with HTTPS) and which model provider to use; flags answer the same
questions for scripts. Re-running keeps secrets and settings and moves the
image tag, and the database password is only generated for a new database.
Also: down, status, logs, token, open, env, upgrade (uv tool installs
upgrade themselves and run `up` again) and uninstall (keeps settings and
data unless --purge). The commands import no Flask, Celery or settings.
The wheel carries deployment/docker-compose-standalone.yaml as
docsgpt/deploy/docker-compose.yaml; the sdist includes the source file.
Outside a checkout the data home was the working directory, so running
`docsgpt api` from another folder silently used different settings and data.
It is now ~/.docsgpt/server (/opt/docsgpt for root on Linux); DOCSGPT_HOME
and a checkout still take precedence. The API and worker commands create
the home and point out a .env left in the working directory.
The backend image builds the web UI with scripts/build_frontend.sh and
serves it through docsgpt/ui.py, so the standalone Compose file drops the
frontend container. UI and API share port 7091, published on 127.0.0.1
unless DOCSGPT_BIND says otherwise. POSTGRES_PASSWORD is configurable, and
an optional https profile puts Caddy in front of a public domain.
docker-image-verify.yml starts the standalone stack on the image it built
and checks the API, the UI, /config.js and a client-side route on one port.
- Ship tiktoken's cl100k_base inside the package and build the encoding
from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
temp dir, and read tokenizer.json and repo metadata from that cache, so
a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
the endpoints return 404, audio files fail to ingest with a clear
message, /api/config reports tts_available/stt_available, and the UI
hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
The connector OAuth popup posted the session token to window.opener with
a '*' target origin, so any page that opened the popup received it. An
attacker with an account on a multi-user deployment could start a flow for
their own pending session, get a victim to finish the provider consent, and
receive a token backed by the victim's Drive/SharePoint/Confluence tokens.
- Post popup results only to allowed frontend origins: the callback origin,
OIDC_FRONTEND_URL, the new CONNECTOR_ALLOWED_ORIGINS, and localhost:5173
when the callback runs on a loopback host.
- Render the success page from the callback itself so the token never
appears in a URL; callback-status ignores session_token/user_email params.
- ConnectorAuth accepts messages only from the popup it opened, on the
callback origin reported by /api/connectors/auth.
- /api/connectors/disconnect requires auth and only deletes the caller's
session.
- /api/connectors/sync and /api/remote reject session tokens the caller
does not own.
Fixes#2766
A deflate-compressed TIFF under 1 MB can declare 144 million pixels and
take 1.2 GB to convert to PNG, and Pillow only warns below 179 million.
Read the dimensions from the header and refuse images over 40 million
pixels before any pixel data is decoded. Pillow's DecompressionBombError
is now raised as DocumentParseError, so the upload fails once instead of
being retried.
- Per-user SSE cap uses per-connection leases in a sorted set instead of
a shared INCR/DECR counter with a TTL. A stream that outlived the TTL
could let the counter expire, then decrement another stream's slot or
drive the count negative. Leases refresh while a stream sends frames
and age out when a stream dies without cleanup.
- Shielded cleanup is bounded per step: stream close and on_close in
ClosingStreamingResponse, unsubscribe and close in AsyncTopic, lease
release, and the replay-budget check, so a dead Redis connection can't
hold a request or a graceful shutdown.
- Artifact downloads send an ASCII filename plus an RFC 5987 filename*
for non-ASCII names, build headers before opening the file, and close
disk handles off the event loop.
- Type hints and docstrings on the new helpers.
Move GET /api/events, the remote-device command stream and artifact
downloads from Flask to Starlette routes mounted ahead of the Flask
catch-all. On Flask each held an a2wsgi threadpool slot for as long as
its response stayed open, and because uvicorn drops writes after a
client disconnects, a closed tab never released it.
- asgi_auth: one JWT/OIDC gate for Starlette routes; the chat reconnect
reader uses it too
- ClosingStreamingResponse closes the body iterator and releases the
SSE slot or file handle even when the client leaves before the first
frame
- AsyncTopic liveness probe replaces the sync client's socket_timeout
guard against half-open pub/sub sockets
- ASYNC_REDIS_MAX_CONNECTIONS sizes the async Redis pool; every open
stream holds a connection and redis-py defaults to 100
The version input reached checkout as a bare ref, so a manual run given a
branch name would have built that branch and published it as an image tag,
`latest` included. It also meant a version that happens to match a branch name
would silently resolve to the branch rather than the tag.
Resolve it under refs/tags/ instead. Plain versions from backend-release keep
working, and anything that is not a tag fails at checkout.
A manual run takes the version as an input but the checkout had no ref, so it
built whatever the run was dispatched from. Dispatching off main to rebuild an
older release would have published current main under that release's tags, and
moved `latest` to it.
Check out the version instead, falling back to the run's own ref when there is
no version, which is the push that publishes `develop`. The release and call
paths already checked out the right commit; this makes them explicit and fails
loudly on a version with no tag rather than mislabelling an image. The compose
file attached to the release now comes from the tag too.
None of the three image workflows could be started manually, so when 0.20.0
tagged before the release chain knew about the frontend and sandbox images,
there was no way to build those images for the tag short of recreating the
release. Add a workflow_dispatch trigger taking the version to build.
A manual run also gets a `move_latest` switch, so rebuilding an older tag does
not drag `latest` backwards. The rule for the moving tag now lives in the
manifest step's shell rather than in a nested expression: `latest` follows the
build unless the release is a prerelease, the run asked it not to, or it is the
rolling `develop` build.
0.20.0 failed to upload with "Certificate's Build Config URI ... does not
match expected Trusted Publisher". Authentication was fine: trusted publishing
matches the called workflow, pypi-publish.yml, which is what PyPI is
configured with. The PEP 740 attestation the publish action attaches by
default records the workflow that STARTED the run instead, backend-release.yml,
and PyPI verifies that against the same publisher entry. The two claims can
never agree while the publish runs as a called workflow, and no publisher
configuration satisfies both.
Attest only when this file is the entry point, which covers the manual
dispatch and a release published by hand. The release chain uploads unattested
rather than failing.
The earlier rehearsal went to TestPyPI and passed, so this only appeared on the
real index.