- check_usage treats a request through a keyless (draft) agent as agent
traffic, matching how its usage rows are bucketed and the headless rule
- dashboard edits carry the stored enabled flag instead of re-enabling the
policy; disabled policies are labelled in the Quotas tab and the editor
- quota 429s send x-should-retry: false so OpenAI SDK clients do not retry
a refusal that cannot succeed before the reset
- cached-input and cache-write rates for Anthropic, OpenRouter and Groq
gpt-oss-120b; refresh OpenRouter deepseek-v3.2 list prices
- UsageQuota reuses usagePercent; docs note that a user override needs an
existing user
Checked against OpenAI's pricing page: gpt-5.5 at $5 / $30 (cached $0.50)
was already right. The mini and nano models declared no cached rate, so
cached prompt tokens were billed at the full input rate.
- A tool continuation refused for usage now releases the resume claim it
took; before, retries got a 409 until the stale claim was reverted.
- Agent traffic is any row with an agent key or an agent id, so keyless
agents and workflow nodes count toward the agent bucket, not direct.
- The user quota modal discards responses for a previously opened user.
- The usage meter shows every limited bucket, not only 'all'.
- Restore the class separator on the analytics stat card that a formatter
run removed, and align the OpenRouter DeepSeek description with its rates.
Validation problems are returned as values rather than raised and echoed
with str(exc), and a huge integer limit is rejected as out of range instead
of overflowing. Tests no longer call mutating endpoints inside asserts.
The unpriced-model notice asked the live registry whether a model has a
price, so a priced model whose provider was later disabled showed up as
unpriced. It now lists models whose calls this period were all recorded at
$0. Token limits are serialized as integers.
How the instance default, team allowances and user overrides resolve
(including users in several teams), the quota window, who is charged for
agent traffic, how cost budgets price models and what happens to unpriced
ones, and the admin and user API.
The admin dashboard gets a Quotas tab for the instance default, team
allowances and user overrides, with a notice listing models that cost
limits cannot see. A Quota action on the Users tab shows a user's effective
limits, the layer each comes from and their usage, next to the editor for
their override. Each budget is either not set at that layer, a limit, or
unlimited.
Users with a quota see a usage meter with the reset time on the Analytics
settings page, and a refused chat request shows the used amount, the limit
and the reset time in the user's language.
Admins read and set the instance default, team allowances and user
overrides under /api/admin/quotas. A user's endpoint also returns the
limits those layers resolve to, the layer each came from and the usage
against them. The overview lists catalog models used this period that have
no price, since a cost limit cannot see them. Every write is audited.
GET /api/user/quota gives a user their own limited buckets, usage and reset
time without naming the policies behind them; any valid token may call it.
check_usage now checks the billable user's quota on every request, before
the per-agent 24h limits, which keep applying to traffic through an agent.
Until now a request without an agent key skipped every limit. A refusal is
a 429 with Retry-After and a body naming the budget, usage, limit, the
layer the limit came from and when it resets.
Headless runs check the agent owner's quota before starting. A refused
scheduled run is recorded as budget_exceeded; a refused webhook run returns
a quota_exceeded result instead of raising, so Celery does not retry it.
docsgpt/quotas resolves a user's effective token and cost limits from the
policy rows that apply to them: their own override, then the most generous
allowance among their teams, then the instance default, then any defaults a
registered provider supplies. Each budget resolves on its own, and a team
counts once however many memberships the user holds in it.
QuotaService compares those limits with the user's token_usage totals over
the current QUOTA_PERIOD window (calendar-aligned, UTC, computed at read
time). It fails open, and skips the usage query for unlimited users.
sum_tokens_in_range now skips the scheduler's per-run rollup rows, whose
tokens are already on the run's per-call rows, so the per-agent 24h limit
and the admin total no longer count scheduled spend twice.
Workflow node LLMs carry the workflow agent's id, so their usage rows are
attributed to the agent instead of landing with a user id only.
usage_totals returns a user's tokens and cost since a window start, split
by interactive and agent-key traffic.
Migration 0033 adds quota_policies (instance default, team per-member
allowance, user override; a token budget and a USD budget per row) and
token_usage.cost. The column is added IF NOT EXISTS so a database that
already carries it upgrades cleanly.
Every usage row now records the call's USD cost from the model catalog;
bring-your-own models are recorded at $0.
Rename the unused *_cost_per_token capability fields to USD per 1M tokens,
add prompt-cache read/write rates, and ship list prices for the hosted
catalogs. The old per-token keys still load, scaled, with a warning.
docsgpt/pricing.py turns a call's token bins into a USD cost. Models with
no declared rate cost $0 unless QUOTA_UNPRICED_RATE_PER_MILLION is set.
Review asked whether two document rows with the same text and metadata should
stay two pages. They should not: the caller gets four pages to hand a model,
and a crawl that ingested the same text twice would spend two of them on it.
The rows differ only by an id the model never sees. Pinned either way now.
The tool now asks graphrag_available(), which wants pgvector as well as the
flag. One case set only the flag and passed locally off a dev .env, then
failed on CI's faiss default. An autouse fixture sets it for the module, so a
case that forgets fails for its own reason; the two about a different vector
store override it themselves.
``canonical_name`` answers "" for a punctuation-only name, which callers are
meant to read as "no entity" -- ``_resolve_endpoint`` already does. Entity
extraction did not, and nodes merge on that key, so every such entity in a
source collapsed onto one shared node that belonged to none of them.
The subject flag was in the GROUP BY, so a chunk two matching entities link --
one the page is about, one merely mentioned in it -- came back as two
identical pages and spent the caller's page budget twice on the same text.
It is aggregated with bool_or now, which is what the ordering wanted anyway.
Covered by a live test against a pgvector-shaped documents table: the graph
tables alone cannot answer this query, so nothing exercised it before.
Three faults in the graph tool, all on the agent's path:
The store was cached on the tool, and the executor caches the tool for the
whole agent run -- so one pgvector pooled connection stayed checked out across
every LLM round trip of that run, minutes at a time, and enough concurrent
runs exhaust the pool. GraphRAGRetriever releases its store before falling
back for this reason. The tool now releases it at the end of each action.
It gated on GRAPHRAG_ENABLED where everything else asks graphrag_available(),
which also requires the pgvector store. Under any other vector store the graph
tables are not the ones the sources were ingested into, but the tool was still
offered and still queried Postgres.
Pages were labelled by hand rather than through labels_from_metadata, which
exists so citation labels match across retrievers. A page read by the tool and
the same chunk retrieved by internal_search are one document, and citations
key on (source, title) -- so the research agent gave that document two
citation numbers. The recorded doc also keeps the full chunk text now, so it
dedupes against the retriever's copy; only what the model reads is truncated.
Only a raise routed a source to the ClassicRAG fallback, and every graph read
logs its own failure and returns empty. So a query that broke, a half-built
graph and a walk that genuinely found nothing were indistinguishable, and each
made the source contribute nothing at all to the answer -- no fallback, no
vector blend, which is skipped by the same early return.
A graph source that produces no documents now joins the classic batch, exactly
as a source with no graph already does.
The passage stage also rescanned every node's chunk list once per candidate.
It inverts the links once instead: 7.3 ms to 0.16 ms on a 400-node subgraph at
the candidate cap, with identical output.
The three per-source graph options are read from the retrieval config the
Dispatcher hands over, and it only hands one over for a source it considers
overridden -- which it decided from chunks, score_threshold, rephrase_query
and prescreen alone. A source that changed only its graph options was not
"overridden", so nothing was carried and every graph source ran the defaults:
the UI toggles did nothing at all.
They count as an override now, for graphrag sources only. They mean nothing to
any other retriever, and an override also hands the source its own chunk
budget, which a classic source must not pick up from a graph setting.
Two builds of one source can overlap: the extraction lease is keyed by the
source's updated_at, and enabling a graph updates the source before it
dispatches, so a rebuild started while the last build runs gets a new key and
a lease of its own. Both builds could then pass a chunk's "done" check before
either committed and apply it twice -- doc_freq bumped twice, reproduced with
two live writers. A reset could also land in the middle of a chunk.
apply_chunk and delete_by_source now take a transaction-scoped advisory lock
keyed by the source before touching a row, as the schema bootstrap already
does for DDL. A single build's writes were already serial, so it loses
nothing; overlapping builds take turns chunk by chunk, and the second sees the
first's "done" row and returns (0, 0).
apply_chunk also still defaulted with `rel.get("weight") or 1.0`, turning an
explicit zero into a full-strength edge -- the conversion 3f774d81 removed from
add_edge and the ranker but missed here. Only a missing weight defaults now.
Four graph-store queries formatted table and column names into the SQL
string: get_chunk_texts and delete_by_source (Bandit B608, alerts #582/#583
on main) and entity_pages/chunk_similarities from this branch (#661/#662,
dismissed). The names were validated by _safe_identifier, so none was
injectable, but each query was still a string built at runtime.
They are now fixed statements composed with psycopg.sql: identifiers go in as
sql.Identifier, values stay bound. Identifiers are lower-cased before quoting,
because PGVectorStore writes the same names unquoted and Postgres folds those
to lower case -- a quoted mixed-case name would address a different table.
Bandit reports nothing for docsgpt/graphrag now; the queries return the same
rows against a real graph as before.
The lifecycle tests imported docsgpt.celery_init as a module next to the
file's from-imports, which code scanning flags. Patch the flag by path and send
the signals from celery.signals -- the objects a worker actually fires.
CI has no pgvector, so every live graph test skips there and the queries this
branch added ran in no CI job at all. These pin what holds without a
database: the four new read queries bind every value (the entity name comes
from an LLM tool call) and map their rows; empty input runs no query; a failed
query returns nothing and releases its connection. The graph tool's plumbing,
the sources_have_graph gate and the hybrid path's vector ranking get the same.
Also drops a redundant chained comparison flagged by code scanning.
The graph tool records every page read_entity_pages returns in
retrieved_docs, but agents only ever collected internal_search's. A turn
answered from graph pages therefore emitted no sources: on the multi-hop e2e
run, a question answered from two read_entity_pages calls came back with none.
_search_tool_docs gathers both tools' documents the way the executor caches
them, and both the classic/agentic collector and the research agent's
per-step citations use it. Checked against the real executor and a real
graph: a graph-only turn now cites the pages it read.
in_worker() read task_join_will_block, which eventlet and gevent leave unset,
and the task's current_worker_task, which they scope to one greenlet -- so a
greenlet a task spawned in those pools still took the web-process branch and
dispatched to its own worker, where it could queue behind its parent and time
out.
The worker's own startup now records it: worker_init fires in every worker's
main process before the pool starts (where solo, threads, eventlet and gevent
run tasks, and what prefork children fork from), and worker_process_init in
each prefork child. worker_ready would be too late -- prefork children are
forked before it fires. The two existing checks stay for anything that runs
tasks without that startup.
Provider-reported usage is kept on the LLM instance (_last_usage) and claimed
by whichever call finishes next. With GRAPHRAG_EXTRACTION_WORKERS > 1 (the
default is 8) every extraction thread shared one instance, so a call could
claim another call's provider counts while its own fell back to the estimate:
token_usage rows, and the cost they bill, could be attributed to the wrong
call and summed wrong.
Each pool thread now builds its own extraction LLM on first use. The calling
thread's instance is still built up front, so a misconfigured model fails the
run before any chunk is touched.
canonical_name dropped "es" from every -ches/-ses/-zes plural, so "caches"
became "cach" while "cache" stayed "cache" -- the singular and plural landed on
two nodes, which is the split the function exists to prevent. The same rule
split the words documentation uses most: databases/database, responses/
response, releases/release, sizes/size.
An "-es" plural cannot say whether it is cache + "s" or batch + "es", so
instead of guessing, both sides now meet at the stem: the singular endings
-che/-she/-se/-ze/-xe drop their "e" the way the plurals drop "es", and a
singular's -ie folds to -y as -ies already did (cookie/cookies). The key is a
merge key that is never shown, so it only has to agree, not be a word. alias,
canvas, atlas and bias join the words that only look plural.
Found by the naming tests CI was missing: the module had no direct tests.
A chunk whose extraction failed was marked failed and the build moved on. The
checkpoint treats failed chunks as pending, but nothing ever reran the build,
so one transient error -- a provider hiccup, a single response that did not
parse -- left a permanent hole in the graph until someone rebuilt the whole
source.
Failed chunks now get one more attempt after the rest of the build, so a burst
of rate limiting has time to pass. Extraction errors, unparseable responses and
failed writes are all retried; a chunk is marked failed only when its retry
fails too, so it costs at most two calls. The chunks given up on are logged by
id, and progress still ends at the total.
Celery records the executing task on the thread that runs it, so a thread that
task starts sees none. The embeddings client and read_document both decided
"am I in a worker?" from that alone, and from any other thread took the
web-process branch: dispatch to the worker they were running in and block on
the result. Celery refuses that get() ("Never call result.get() within a
task!"), so the embed failed and latched the 30s dispatch cooldown for every
caller after it; with joins allowed, read_document would instead wait on a
parsing queue only its own busy process serves.
Threads inside tasks are not hypothetical: per-source retrieval fans out to a
pool, so a scheduled or webhook agent searching several sources embedded from
pool threads. Graph extraction did too, which failed every chunk of a build.
in_worker() in celery_init answers for the whole process. Celery's
task_join_will_block is process-wide and set for every blocking pool (prefork,
solo, threads) -- exactly the condition under which dispatch-and-wait goes
wrong; eventlet/gevent leave it unset, so the task's own thread still counts
through current_worker_task. Verified with real workers on each blocking pool:
from a thread a task started, the old check dispatched and hit the error, the
new one embedded locally.
Exposes the three per-source graph options in the source's retrieval settings,
shown only when the retriever is graphrag: where the walk starts (entities or
relationships), whether passages join the walk, and whether vector hits are
blended in. Defaults match the backend's measured-best configuration and are
filled in for sources saved before the options existed. A note points at the
"search tool" exposure, which is what offers the graph to an agent.
A one-shot graph ranking diffuses over the whole neighbourhood; a question
whose answer sits two hops away is better served by following the edges. The
graph_search tool gives an agent search_entities, get_relationships and
read_entity_pages over its graph sources. On a multi-hop corpus where the
bridging entity is never named, answers went from 0/10 with classic vector
retrieval to 10/10 with the tool, end to end through /stream.
The tool has no setting of its own. It is offered exactly where the agent can
already search: agentic and research agents always, a classic agent only for
sources exposed as a search tool. A graph source left at prefetch is used for
ranking only.
Tests now pin GRAPHRAG_ENABLED to its shipped default, as CI has: with a dev
.env enabling it, every agent test's graph check read the developer's real
database and left a pool to it behind.
Graph retrieval tied plain vector search at best and never beat it. Measured
across five corpora, the bottleneck was seeding, not the graph: the walk
started from nodes whose embeddings were computed from bare entity names, and
a whole question shares almost nothing with a name like "Quill".
Extraction now embeds each node from "name (type): description" and each
relationship as the fact it asserts ("Alder streams_to Quill: ..."), stored on
a new nullable graph_edges.fact_embedding column that ensure_vector_schema adds
in place. Entity names are canonicalised (case, punctuation, word breaks and a
cautious plural) so "VECTOR_STORE" and "vector stores" land on one node. Extraction calls run
concurrently (GRAPHRAG_EXTRACTION_WORKERS, default 8) while embedding and graph
writes stay serial on the task thread, so ordering and idempotency are
unchanged; that measured 8.4x faster with identical output.
Retrieval gains per-source options, stored under retrieval.graph and read live
at query time:
- seed_strategy: start from matching entities (default) or matching
relationships, which can reach an entity the question never names;
- passage_nodes (on): walk the source's passages alongside entities, with
PageRank damping 0.5 instead of 0.85;
- blend_vector (on): fuse the graph ranking with the source's vector ranking
by reciprocal rank.
The defaults are the measured-best configuration. Through GraphRAGRetriever,
the new seeding moved recall@4 from 0.41 to 0.68 on a multi-hop corpus and
from 0.50 to 1.00 on the docs corpus, and regressed none of the corpora
measured. Existing graphs keep name-only embeddings until rebuilt.
Returning a deferred marker traded one wrong signal for a worse one. A
redelivery reuses the original task id — Context.as_execution_options carries
task_id into the retry — so the duplicate's return marked the very id the
client polls as SUCCESS. /api/task_status reports celery's state verbatim and
the UI maps SUCCESS to "done", so the GraphRAG enable modal would announce a
finished build, rendered from a payload with no counts, while the run holding
the lease was still extracting.
Raise Ignore instead: celery records no state for the duplicate, so the task
id keeps whatever the holder sets and the poller keeps waiting. The autoretry
wrapper re-raises Ignore ahead of autoretry_for, so the wider
autoretry_for=(Exception,) on these tasks cannot turn it back into a retry.
Three findings from review, all on code this branch introduced.
apply_chunk was not replay-safe. commit() can report connection loss *after*
Postgres committed, and the reconnect retry then replays the write: _upsert_node
bumps doc_freq a second time and _add_edge inserts another row, since
graph_edges has no uniqueness constraint for a logical edge. The chunk's
graph_ingest_progress row is now written in the same transaction as the rows it
describes, and a replay that finds it already "done" returns (0, 0) without
touching the graph. Extraction drops its separate mark_chunk("done"): the
checkpoint and the graph can no longer disagree.
count_nodes swallows every query failure and answers 0, so extraction's
"fall back to the write count" handler could never run — a failed count after a
successful build reported an empty graph. count_nodes grows a strict mode that
re-raises; retrieval keeps the swallow, which is what routes a source to
ClassicRAG.
A zero edge weight was read as a full-strength link: `or 1.0` rewrote an
explicit 0 before the <= 0 filter. Only missing and null weights default now.
The same coercion sat in _ppr_scores, where it would have kept the ranker's
rule unreachable from the product path, so it is fixed there too.
Review feedback on _write_with_reconnect: an explicit return inside the loop
plus an implicit fall-through past it reads as a path that returns None. The
retry is exactly two attempts, so say so — first attempt, reconnect on
connection loss, second and final attempt — and every path now returns or
raises.
Also narrows the transitions hint to the (node, weight) tuples it holds, and
logs at debug when the power iteration hits its iteration cap instead of
returning the last iterate with no trace.
A task that runs longer than the broker's visibility timeout is redelivered
while its first run is still going. The idempotency lease correctly stops the
duplicate from doing the work, but the duplicate then re-queued itself once
per LEASE_TTL until celery ran out of retries and raised
MaxRetriesExceededError — so a perfectly healthy long task (a large graph
extraction is the one that found this) reported a task failure, with a
traceback, while the real run was still making progress next to it.
Catch the exhaustion and return a "deferred" result instead. A normal
deferral still re-queues: only the give-up path changes, and the lease
holder's dedup row is left untouched so its own completion still records.
Three ways a chunk disappeared from a graph with no way to tell:
A build checks out one connection and then spends minutes per chunk waiting
on the model, so the connection idles long enough for the server or a pooler
to drop it. The pool only validates a connection when it hands one out, and
this one was handed out at the start of the build, so the next write raised
"the connection is lost", the chunk was marked failed, and the build carried
on a chunk short. apply_chunk and mark_chunk now reconnect and retry once;
every statement they run is an idempotent upsert, so a replay cannot
double-write. Only connection loss retries — a bad statement still surfaces.
An unparseable model response marked the chunk failed and logged nothing at
all, so failed_chunks was the only evidence and it named no chunk. Both
failure modes now log the chunk id.
The summary's node count summed per-chunk upserts, so an entity appearing in
ten chunks counted ten times: it reported writes, not graph size. It now
reports the distinct node count, falling back to the write count only if the
count query fails.
_build_extraction_llm passed settings.LLM_PROVIDER with the resolved
extraction model id. Those two disagree in any deployment that leaves the
provider at its default: the model id comes from GRAPHRAG_EXTRACTION_MODEL
or LLM_NAME, while the provider stays "docsgpt" — the hosted public
endpoint, which does not serve it. The request is rejected, the shared
fallback answers instead, and the graph gets built by a different model
than the one configured, with nothing in the summary saying so.
Resolve the provider from the model registry (owner-scoped, so per-user
BYOM ids resolve too), fall back to settings.LLM_PROVIDER only when the
model is unknown, and take the API key for the provider actually dispatched
to rather than the generic settings.API_KEY. The effective provider is
logged, since a silent swap was the whole failure mode.
networkx.pagerank delegates to a scipy implementation, and scipy is not a
DocsGPT dependency — it only arrives transitively through the optional
docling extra. In a default install every graph retrieval raised
ModuleNotFoundError inside _ppr_scores, hit the per-source except, and
degraded to ClassicRAG: the graph was built and paid for, then never used,
with one ERROR line per source per query as the only signal.
Rank with a local power iteration over the same row-normalized transition
matrix: undirected edges normalized per endpoint, dangling nodes
redistributed along the restart vector, and the restart vector normalized
across the nodes the subgraph actually holds so seed mass cannot leak.
Parity with networkx is asserted while scipy happens to be installed in the
test env, and the retrieval path is exercised with the import blocked.
_redis_urls raises before any check runs and cli.main prints what it raises, so
every one of its five messages interpolated the URL it was given — password and
all — straight to stderr. They render it through _endpoint now.
_endpoint itself then echoed the whole path, so an over-long database number
came back at full length: 5132 characters of error for one bad setting. It caps
the path it renders, which protects every caller, since that string is written
into terminals, CI logs and error messages rather than reused as a URL.
CodeQL flags "host" in message as incomplete URL sanitisation. I removed that
shape from two assertions and introduced a third in the same commit; this takes
it out of the provider test and the remaining Redis one, comparing with what
_endpoint produced instead, which is the contract those tests actually mean.
- doctor printed OPENAI_BASE_URL raw, the third place a credential-bearing URL
reached the terminal; it goes through _endpoint like the others.
- dev.run left its output readers unjoined, so a child's last lines could be
lost on exit. The threads are kept and joined during teardown.
- logs read each file and then reopened it to follow, so anything written in
between appeared in neither. One handle now serves both.
- dev took --port straight from argparse: 0 would have served on an ephemeral
port while printing 0, and oversized values reach socket.bind. It goes
through _port_number first.
- doctor --redis-url overrode only the broker, so a stale result backend or
cache was still pinged and the flag looked broken. All three endpoints now
come from the URL given.
CodeQL flags `"host:port" in message` as incomplete URL sanitisation. It is a
message rather than a URL being authorised, so it is not a vulnerability, but
the check was red and comparing against the sanitiser's own output asserts the
real contract. The signal handler now says why it swallows: the child is
already gone, and shutdown must not fail on what it is cleaning up.
env.py sets no version_table_schema, so the table follows search_path. Pinning
public. in doctor's queries made them agree with each other and disagree with
alembic: on an install using another schema it reported no schema at all and
sent the user to migrate an already-migrated database. Both queries resolve the
table the same way alembic does now.
Both preflights passed for `dev --mock-llm --port 8090`: the port was free, and
then the mock and the API were each handed it. One child could not bind, and the
API was pointed at that port as its model server while trying to listen on it.
The mock starts first and the API and worker are pointed at it, but only the
API port was preflighted. A busy 8090 therefore showed up as a child exiting
once the rest were running, or as the API talking to whatever else was on that
port. --ui is deliberately left alone: vite.config.ts sets no strictPort, so
Vite moves to the next free port rather than failing.
urlsplit accepts redis://host:notaport/0; only parts.port raises, and it raises
on access rather than at split time, so the ValueError fell outside the try.
_check_redis catches the client error and then formats it through _endpoint, so
doctor ended with a traceback from inside its own error path.
Running outside a checkout and starting on a port something else holds are both
refusals a developer will meet, and neither had a test. The second also asserts
that nothing is spawned when the port is taken, which is the part that matters:
the guard runs before any child process exists.
Sanitising the URL in the message was not enough: the client's own error text
went into the detail too, and both psycopg and redis-py quote the URL they were
given. The reason is kept, the endpoint is kept, and the URL, username and
password are taken out of it.
The Redis test raised a generic error, so it asserted the password was absent
without ever exercising the path that leaked. Both tests now raise what the
clients actually raise.
The Redis check put the whole URL in its failure message. A managed Redis URL
carries user:password@host, so a failed ping printed the password to the
terminal and into any log or issue the output was pasted into. It names
scheme://host:port/db now, and falls back to naming no URL when the value
cannot be parsed at all.
The Postgres check asked to_regclass about public.alembic_version and then read
version_num through search_path, so another schema could answer with a
different revision, or the query could fail, and doctor would send you to run
migrations against a database that is already fine.
_port_number was written so a hand-edited .env could not reach int() raw, and
then doctor did exactly that: a nonnumeric port ended the command with a
traceback rather than the message, in the one command whose job is to explain a
broken setup.
The checks were mocked wholesale, so the branching that produces each diagnosis
had never run: a database with no schema yet, one behind this version, one that
refuses the connection, and which of the three Redis URLs failed. Each of those
is the sentence a developer reads when something is wrong, so each is pinned.
_migration_head is tested against the packaged alembic.ini itself: it needs no
database, and it is the path resolution that breaks silently when files move.
`docsgpt up --native` installs services meant to outlive the shell. Development
wants the opposite, and until now it meant three terminals from the guide:
uvicorn, celery, and vite.
`docsgpt dev` runs this checkout's API and worker as children of one terminal,
both restarting when a file is saved, their output interleaved and labelled, and
Ctrl-C stopping them together. `--ui` adds the Vite dev server, `--mock-llm`
runs the bundled mock model so no API key is needed, and `--no-worker` leaves
the worker to your editor's debugger. Celery has no reloader of its own, so the
worker is wrapped in watchfiles when it is installed, and runs plain when it is
not.
Alongside it, the commands a dev loop keeps reaching for:
- `docsgpt doctor` checks what usually breaks a new setup: PostgreSQL answering
and its schema matching this version, Redis answering, a model provider being
configured, and the port being free.
- `docsgpt restart [api|worker]` bounces services without rewriting settings or
rerunning migrations, which `down` plus `up` did.
- `docsgpt logs -f` follows a native install instead of telling you to run
`tail -f` yourself.
- `docsgpt env set` applies itself to a running native install rather than
asking you to run `docsgpt up` again to change one value.
Two bugs found on the way, both older than this change:
- `docsgpt api --reload` watched the working directory, which in a checkout is
178,425 files: .venv, node_modules, and the indexes/ and inputs/ the app
writes to while ingesting, so the server restarted itself mid-request. It
watches the package now — 1,217 files.
- The VS Code "Flask Debugger" ran `flask run`, which serves only the WSGI app:
/mcp, the SSE streams and artifact downloads 404 under it. The guide warned
about this in prose while the debug config did it anyway. It runs uvicorn on
the ASGI app now, like production.
parts.port returns 0 rather than raising, since 0 is inside the range it checks,
so the URL reached .env and the worker and cache had nothing to connect to.
urlsplit accepts an authority such as localhost:notaport or localhost:65536 and
urlunsplit rebuilds it verbatim; only parts.port raises, and nothing read it. The
unusable value reached .env, where the worker picked it up and failed to start
its broker, while `up` reported success because it waits only on the API health
endpoint.
The ASCII-digit check accepts any length, but since 3.11 Python refuses to
convert a digit string past its conversion limit, so a long one raised
ValueError straight through the CLI instead of the message every other
unusable URL gets.
The exemption asked whether any service of the install was running, so moving an
install onto a different port that something else held would pass the check and
then fail to bind, with the health poll answered by whatever owned that port —
the false success the check exists to prevent.
install.json carries the API port now, and a busy port is allowed only when it
is that port and the API service is running.
The Docker-stack check ran after logs/ was created, .env written and the
migrations applied, so a conflict left a migrated database and partial files
behind with no install.json — exactly the directory that down and uninstall
then refuse. It runs immediately after the port is resolved now.
Alongside it, `up --native` refuses a port it cannot bind. Neither service
manager confirms that the API bound, and /api/health carries no installation
identity, so a second install on the same port would have been answered by the
first and reported success while its own API was dead. An install re-running on
its own port is the exception, since its services are what hold it.
From the outside-diff findings on #2800:
- Service names are derived from the install directory. A service manager has
one namespace per user, so two installs in different --dir directories wrote
over each other's units and down, status and uninstall acted on whichever was
written last. The default install keeps the readable names; another directory
gets a digest suffix.
- `up --native` refuses when a Docker stack in another directory publishes the
same port: its API would answer the health check while these services failed
to bind. The check degrades quietly when Docker is absent, which is exactly
the machine a native install targets.
- A Redis database path is required to be ASCII digits: str.isdigit() is true
for characters int() then refuses.
- DOCSGPT_PORT from a hand-edited .env is validated before conversion, and the
error names where the bad value came from.
- Percent signs are doubled in systemd values, arguments and log paths, since
systemd expands specifiers in all of them.
urlsplit raises ValueError on input such as redis://[::1 , which nothing
converted, so a typo left native setup with a traceback rather than the message
every other unusable URL gets.
The three URLs were built by string surgery, so anything after the database
number was mangled rather than kept: rediss://host:6380/0?ssl_cert_reqs=required
came out as .../0?ssl_cert_reqs=required/0, and a URL carrying a query but no
database had /0 appended after the query. TLS and managed Redis endpoints
usually carry exactly those parameters.
The URL is split properly now, the three databases go in the path, and scheme,
credentials, host, query and fragment are preserved. A URL that cannot be
numbered this way — one that is not redis:// or rediss://, or that has
something other than a number where the database goes — is refused with a
message instead of being turned into something that merely looks like a URL.
Quoting cannot carry a newline into a unit file or a plist: the line ends and
whatever follows becomes another directive. Every value bound for a service
file — the working directory, the log path, environment names and values, and
the command arguments — is checked before any of it is rendered, on both
launchd and systemd.
`up --native` took --expose, --domain and --docling and did nothing with them.
Asking for network exposure and silently getting a loopback-only install, or
asking for docling and getting an install without it, is worse than being told.
Each now says what to do instead: a reverse proxy or the Docker stack for
exposure, and the docling extra for the parser engine. --expose local and
--no-docling already describe native mode, so they stay silent.
Both accepted a native install and then drove `docker compose` in a directory
with no compose file, so the user got "no configuration file provided" rather
than an explanation. They now refuse with what to do instead, and restore
refuses before it reads the archive or stops anything.
`open` built its address from stack.url, which honours DOCSGPT_BIND, while the
native units always bind 127.0.0.1: with a LAN bind it handed the browser an
address nothing was listening on. status had the same mismatch and was fixed
with it; both now go through one helper so they cannot drift apart again.
From the outside-diff findings on #2800:
- install.json is written before the services are installed and started. A
service that fails to start used to leave units behind in a directory that
status, down and uninstall no longer recognised as a native install, so
nothing could clean them up.
- systemd stop and removal propagate failures: `down` reporting success while
the unit still runs, or `uninstall` dropping the unit file and the record
while systemd still runs the service, is worse than an error. Removing a unit
that is already gone stays harmless.
- WorkingDirectory and each Environment value are quoted and escaped for
systemd. `--dir` takes a free-form path, and one with a space in it is not
hypothetical: this checkout lives in one.
- Native status checks and prints http://localhost:<port>, which is what the
units bind. With a LAN DOCSGPT_BIND it used to poll an address nothing
listened on and call a healthy install dead.
`upgrade` exec'd the bare name `docsgpt`, so after `python -m docsgpt upgrade`
in a virtualenv without the console script on PATH, os.execv failed with a
traceback. It now uses the same launcher the service units get, which is why
that helper is no longer named for native mode.
Refusing to write the service units when `docsgpt` is not on PATH was wrong. A
package installed in a virtualenv is runnable whether or not its console script
is on PATH, and CI runs pytest as `python -m pytest`, where argv[0] is a module
file: the refusal failed thirteen native tests there.
The launcher now prefers the command on PATH, resolved to an absolute path
since PATH can hold relative entries, then an argv[0] that can be executed, and
otherwise this interpreter with `-m docsgpt`, which works wherever the package
is importable. `python -m docsgpt` became an entrypoint of its own and has a
test that runs it.
From review of #2800:
- systemd `enable --now` starts nothing when the unit is already active, so a
second `up --native` kept the old ExecStart and left the API on its previous
port. start enables and then restarts, as the launchd path already did by
booting the job out first.
- An explicit `home` now wins over XDG_CONFIG_HOME, which is what callers pass
it for.
- The ExecStart program must be a real executable: when `docsgpt` is not on
PATH, sys.argv[0] is accepted only if it can be run, and otherwise the
failure is raised before any unit is written.
- `up --native` over a directory holding a Docker install now refuses and says
how to proceed, instead of starting native services beside containers that
down, status and uninstall would no longer see.
- Docs: without a terminal only --postgres-uri is required, and the Windows
fallback names `docsgpt beat`, which the worker cannot embed there.
SystemdServices was the least covered part of the module and cannot be run on
this machine, so it now has tests for install, start, stop, remove, is_running
and a failing systemctl.
Both shutdown calls sat outside the recovery that undoes them: a `compose down`
that failed partway left the stack down, and a `compose stop` that failed left
the backend and worker stopped. Each now runs inside its own try.
A restore also replaced volumes one at a time, checking each tar as it reached
it, so a damaged third payload was found with the first two already swapped in.
Every declared tar is read through first, and the imports start only once they
all come out whole.
A fixed path under /tmp was both a guess about what the image can write to and
a temp-file smell that Bandit flags. The container makes the directory itself
with mktemp -d and removes it afterwards.
Validating the archive catches a damaged one while DocsGPT is still up, but a
well-formed archive can still hold a corrupt volume tar or a dump statement
psql refuses, and those only surface once the stack is down. The work after
the shutdown now runs inside an error boundary that starts the stack again
before the failure is reported, so a failed restore never leaves the install
stopped.
The image does not run as root, so the staging directory could not be created
at the container root: `mkdir /stage` failed with permission denied and every
restore would have failed. It goes under /tmp now, and the command is built as
one string instead of concatenated pieces inside the argument list.
Checked against a real volume and the published image: a truncated tar fails
and leaves the volume exactly as it was, and a whole one restores it.
From review of #2799:
- The archive is created 0600 rather than at the process umask: it holds the
install's data, and --with-settings puts .env and its secrets in it.
- restore validates everything the manifest declares before the stack is
stopped, so a damaged archive fails while DocsGPT is still running rather
than after `compose down` has taken it away.
- Only the volumes a backup is made of are restored. A hand-made manifest can
no longer point import_volume at postgres_data, whose contents it empties.
- psql runs with ON_ERROR_STOP=on, so a restore that fails halfway cannot
start DocsGPT again and call it a success.
- import_volume unpacks into the container's own filesystem first and clears
the live volume only once the tar has come out whole, so a corrupt one
leaves the volume as it was.
- The backend and the worker stop while the archive is made and start again
even if the dump fails, so the dump and the volume tars describe the same
moment instead of drifting apart as ingestion writes.
Run DocsGPT without Docker: the API and the worker each become a service
on the machine itself, a launchd agent on macOS and a systemd user unit
on Linux, pointed at a PostgreSQL and a Redis that already run.
`docsgpt up --native --postgres-uri ... --redis-url ...` writes the same
.env a Docker install uses, applies the migrations and starts both
services. status, logs, down and uninstall work on a native install the
same way they do on a Docker one, and never touch the database or Redis:
they were the user's to begin with.
One Redis URL covers the broker, the result backend and the cache on
three consecutive databases, starting at the one the URL names, so a
Redis that already holds something else can be shared.
Windows has neither service manager, so native mode refuses it and says
what to do instead.
`docsgpt backup` writes one archive holding a pg_dump of the database, a tar
of each data volume and a manifest of what it came from; `docsgpt restore`
puts it back over an install. The settings file is left out unless
--with-settings asks for it, since it holds the install's secrets, and a
backup taken with a newer DocsGPT is refused without --force.
The volume tars go through the image the install already runs, so a backup
pulls nothing extra, and compose calls can now redirect stdout and stdin so
the dump never passes through this process.
The Windows installer only checked that uv.exe exists after running the uv
installer, so a failed install that left an older uv.exe behind was accepted;
it now fails on a nonzero exit code.
sg runs its command with /bin/sh, which need not be bash, so the handoff
after installing Docker quotes each argument as POSIX single quotes instead
of with bash's printf %q.