100 Commits
Author SHA1 Message Date
Alex 89bd6d6071 Merge pull request #2863 from arc53/fix-remove-connect-more-btn
fix: remove connect more link
2026-09-30 00:03:27 +01:00
Alex 411ac91bfd fix: remove connect more link 2026-09-29 23:59:11 +01:00
Alex ffacf5fbb9 Merge pull request #2843 from arc53-machine/connectors
Connectors: one place to connect accounts for Sources and Tools
2026-09-29 20:18:35 +01:00
Alex 7565aa9c4c Merge pull request #2855 from arc53/roles-revamp
Roles revamp
2026-09-29 13:27:59 +01:00
Alex 335f2af70c Merge pull request #2841 from arc53-machine/default-chunks-6
Default retrieval to 6 chunks instead of 2
2026-09-28 15:19:02 +01:00
Alex 1c19ed8ce2 Merge pull request #2840 from arc53-machine/math-rendering-parser-level
Render chat math in the markdown parser: \( \), \[ \], $x$, mhchem, streaming
2026-09-28 13:50:39 +01:00
Alex aa0ce42fd7 Merge pull request #2839 from arc53-machine/failed-attachment-never-blocks-send
A failed attachment never blocks Send: drop it and send the question
2026-09-28 12:21:30 +01:00
Alex d37e0bd1ad Merge pull request #2834 from arc53/iphone-modal-fixes
Iphone modal fixes
2026-09-27 10:42:34 +01:00
Alex 420b4d33b9 Merge pull request #2833 from arc53/ui-polish 2026-09-26 21:04:02 +01:00
Alex 431a069d4c Merge pull request #2832 from arc53/shadcn-linter
Shadcn linter
2026-09-26 11:52:09 +01:00
Alex 9e0272c33f Merge pull request #2830 from arc53-machine/chore/docs-widget-0.8.0
Update docs widget to docsgpt-react 0.8.0
2026-09-25 10:54:15 +01:00
Alex 7b93e67fea Merge pull request #2829 from arc53/chore/bump-npm-v0.8.0
chore: bump npm libraries to v0.8.0
2026-09-25 10:46:38 +01:00
Alex f0aa6ae275 Merge pull request #2826 from ManishMadan2882/main
(fix:widget/ui): mobile keyboard, back button
2026-09-25 10:36:39 +01:00
Alex a17ebf2fc9 Merge pull request #2828 from arc53-machine/chore/banner-secure-oss-fund
Point the dev banner at the Secure Open Source Fund post
2026-09-24 16:46:55 +01:00
Alex f2ad13be57 Merge pull request #2827 from arc53/fix/continuation-log-context
Bind the log context for tool continuations
2026-09-24 13:54:24 +01:00
Alex cf9a5e9cb8 Fall back to model_id for the logged model
Agents store their model only in model_id, so neither log_activity nor
agent_log_context bound a model key. start_agent_span already reads it.
2026-09-24 13:47:46 +01:00
Alex 9df90ab9df Bind the log context for tool continuations
gen_continuation is not wrapped by @log_activity, so the llm_* log lines
from a resumed turn carried no user, agent, conversation or endpoint. Every
OpenAI-compatible /v1 client-tool round landed anonymous in the logs.

Add agent_log_context, which binds the same keys log_activity binds while
keeping any enclosing activity_id, and wrap gen_continuation in it. Both
now derive the keys from one helper.
2026-09-24 13:26:49 +01:00
Alex 91d670b7bb Merge pull request #2825 from arc53-machine/feat/genai-traces
Execution traces: GenAI OpenTelemetry spans and a trace waterfall in Logs
2026-09-24 00:12:12 +01:00
Alex 8a9f3a75d6 Merge pull request #2823 from arc53-machine/fix-chunk-token-counts
Show token counts on source chunks
2026-09-22 16:45:42 +01:00
Alex cf3ae3cdea Merge pull request #2822 from arc53-machine/feat/admin-activity-and-usage
feat(admin): activity feed, data-plane audit, and spend/latency in usage
2026-09-22 13:32:50 +01:00
Alex 6c5dfbe9a0 Merge pull request #2807 from simpleqt/sq919/e2e-plan-link
docs(e2e): drop the link to the never-committed e2e-plan.md
2026-09-22 11:38:47 +01:00
Alex 0c85d766dd Merge pull request #2820 from arc53/ui-scan
UI refinements
2026-09-22 11:22:43 +01:00
Alex e6563549cf Drop two ui-smoke tests the sidebar navigation rework outdated
Both fail on main as of f5b02ed3, before any of this branch's changes —
verified by running them against main in a clean worktree, where they fail
with byte-identical errors. Neither is testing a broken feature; both hang
their assertion on a DOM landmark that #2819 moved.

C13 waited for AgentPageHeader's `Agent sub-navigation` nav on
/agents/manage/logs/:agentId. That page now renders CurrentSectionHeader +
SectionPills, and WorkflowBuilder is the only remaining caller of
AgentPageHeader, so the landmark never appears. C2 proved a locale switch
had taken effect by looking for the Settings entry as a `link` role, which
60532ec4 moved into the sidebar as something else.

The behaviour underneath both still works — the logs page renders its
panels, and switching locale swaps the strings and persists to
localStorage. Only the selectors are stale, so these could be repaired
rather than dropped if someone wants the coverage back.

Removed insertStubAgent with C13, its only caller, and the two imports
that went with it.
2026-09-22 10:18:36 +01:00
Alex 2f1b728e23 Let a fork's own catalogs live in docsgpt/core/models without failing CI
test_each_provider_has_expected_ids asserts the exact set of model ids per
provider against a snapshot. For the first-party catalogs that is what we
want: the snapshot exists so an upstream id typo fails before it silently
breaks every agent that references the old id.

openai_compatible is different. It is the zero-Python extension point —
the README's "Add an OpenAI-compatible provider" section and
examples/mistral.yaml.example both tell you to drop a YAML next to the
built-ins, and the loader globs the whole directory. So anyone who follows
those instructions gets a permanently red test on their fork, for having
done the documented thing. The _internal suffix carve-out only covers the
one filename we happen to gitignore.

Treat the snapshot as a floor for that provider: the built-in ids must all
still be present, extra ones are somebody else's provider. The rename this
guards against still drops an id, so it still fails, and the assertion
names the missing ids instead of printing two large sets to diff by eye.

Verified by copying examples/mistral.yaml.example into the catalog
directory: red before, green after, and still red when a built-in id is
renamed.
2026-09-22 10:03:12 +01:00
Alex c2e04fa7bf Finish the MCP auth labels, the plural forms and the skeleton's status role
Three things the UI pass started but did not carry all the way.

The MCP server dialog's Basic Authentication branch was fixed to use the
label keys for its labels and the placeholder keys for its placeholders,
but the API Key and Bearer Token branches above it in the same
renderAuthFields were left reading placeholders.apiKey and
placeholders.bearerToken as their <Label>. Both fields showed the
placeholder text twice over: the label above the input read "Your secret
API key" where the input already hinted it. authTypes.apiKey and
authTypes.bearer already held "API Key" and "Bearer Token". (Note that
the Basic branch that prompted all this is unreachable — `basic` is
commented out of the authTypes list, so that form cannot be opened.)

Dropped a dead `t(...) || 'Scopes (comma separated)'` fallback in the same
file while there. t() returns the key string when a key is missing, never
a falsy value, so the right-hand side could not run; the branch would
render the raw key rather than the English it looks like it guards.

The plural fix covered English only. resultSummary, view_more,
actionsFound and needsSetup all took _one/_other in en.json, but de, es
and ru kept a single form, so "1 Chunks abgerufen", "1 fragmentos
recuperados" and "ещё 1 источников" still rendered — the exact bug the
change set out to remove, in every language but the one it was tested in.
Added _one/_other for de and es, and _one/_few/_many/_other for ru, which
needs the three-way split (1 источник, 2 источника, 5 источников).
Translated needsSetup while there; it was still English in all three.
jp, zh and zh-TW have a single plural category, so their base form is
already correct and stays as it is.

Replacing Spinner with a skeleton dropped the role="status" that Spinner
carried, and the skeleton divs announce nothing, so a screen reader got
no signal at all between navigating to Agents and the list appearing.
Put it on SkeletonLoader itself rather than at the call sites, so every
variant gets it. It is sr-only, which is position: absolute, so it never
becomes a flex or grid item in the containers these render into.
2026-09-22 10:02:56 +01:00
Alex 327055b5f2 Merge main into ui-scan
The branch forked before the agents navigation moved into the sidebar, so
main had since rewritten both files this PR touches. Two conflicts.

AgentsList: main split the filters into their own routes and dropped the
FILTER_TABS pill row and the page <h1> for CurrentSectionHeader +
SectionPills, so the page-level skeleton this branch added no longer has
a place to live — it keyed off `activeFilter === 'all'` and gated the
whole page on all four sections having loaded. Resolved to main's
structure, which already loads each section independently, and kept only
the part that was the point: the section's Spinner becomes the agentCards
skeleton. That also drops `isLoading && !isDataLoaded`, which would have
suppressed the indicator entirely on a refetch over cached data.

Keeping per-section loading is the better behaviour anyway. The page-level
loader hid the section headers and the New Folder / Import / New Agent
buttons for as long as the slowest of the four endpoints took, and hid the
sections that had already arrived along with them.

AgentLogs: main moved the page onto CurrentSectionHeader + SectionPills
and a max-w-6xl wrapper. Took that, with this branch's removal of the two
page-level spinners — Analytics, GuardrailEvents and Logs each have their
own skeleton, so the outer spinners were the second of the two loaders the
PR set out to remove. Gated on `agentId` alone rather than
`agentId && (loadingAgent || agent)`: the children only ever needed the id,
and the longer condition mounted all three, then unmounted them mid-flight
when the agent fetch failed, leaving the page blank with no error.

The hook-order crash this branch still carried in AgentSection is fixed on
main already (the breadcrumb useMemo sits above the early returns), and the
merge takes that.
2026-09-22 10:02:15 +01:00
Alex f5b02ed374 Merge pull request #2819 from arc53/sidebar-section-nav
Move settings and admin navigation into the sidebar
2026-09-22 00:58:56 +01:00
Alex f50df10311 Guard the second manage-agents handler before its side effects too
The previous commit claimed both handlers, but only the first one
changed: the search text expected `if (...) return;` on a single line
while prettier had already wrapped it across two, so the second
replacement silently matched nothing. A ctrl/cmd/shift-click on the
collapsed-list entry therefore still closed the mobile nav and cleared
the selected agent in the tab it left behind.
2026-09-22 00:54:57 +01:00
Alex 2da6445934 Price the new catalog models and refresh the registry snapshot
main has been red since 295969b: the model catalog gained claude-fable-5,
gpt-5.6-sol and the alibaba and zai providers, but the pieces that track
the catalog were not updated alongside them. Five tests across
test_model_registry_yaml.py and test_pricing.py have been failing since.

- EXPECTED_IDS picks up the two new hosted models and the two new
  openai_compatible ones. The snapshot is deliberate — it exists to
  catch an unintended rename — so it is updated, not relaxed.
- test_everything_set sets DASHSCOPE_API_KEY and ZAI_API_KEY alongside
  the DEEPSEEK one it already set. Every openai_compatible catalog reads
  its own key from the environment, so without them "everything set"
  quietly excluded the two new providers.
- claude-fable-5 and gpt-5.6-sol had no rates, which
  test_hosted_builtin_models_are_priced requires of every hosted model,
  and which token_usage.cost and the quota ceilings both read. Priced on
  the existing ladder: fable-5 above Opus 4.7 at 10/50 per million with
  cached input at a tenth and cache writes at 1.25x, matching the other
  Anthropic entries; sol alongside the gpt-5.5 flagship at 5/30.
2026-09-22 00:42:26 +01:00
Alex f658515ca3 Address review findings on the agents navigation
- Opening a folder dropped the active filter. agentsListPath() always
  built off the bare /agents/manage root, so from /agents/manage/mine it
  navigated to /agents/manage?folder=..., and filterFromPath() then read
  back `all`. A regression from making the filters routable: they used
  to be component state, which a replace-navigation left alone. The
  helper now takes the filter and builds off agentsFilterPath(), which
  also fixes the other half — switching filter while inside a folder
  re-navigates to the filtered path with the folder kept, instead of the
  sync effect pulling it back to the unfiltered URL.

- Both manage-agents handlers cleared the selected agent and closed the
  mobile nav before returning early for a modifier-click, so opening the
  list in a new tab mutated the tab you stayed on. The modifier check
  and preventDefault() now run first.

- The detail breadcrumbs had the name precedence inverted. Tools.tsx
  shows `customName || displayName` throughout, so a renamed tool kept
  its old name in the breadcrumb alone. ToolConfig now matches, and
  RemoteDeviceConfig leads with device?.name as its own heading does.
2026-09-22 00:42:17 +01:00
Alex 448de3c48a Use a lighter icon for Custom Models
Boxes is the densest glyph in the settings nav by some way — nine
elements and 43 path commands against a median of about twelve. Three
cubes with their internal facets inside 20px puts the strokes almost on
top of each other, so at the nav's 1.75 weight it renders as a darker
blob than the icons above and below it, even though the stroke width is
identical. Blocks sits at the median, matching Sources, Tools and Teams,
and stays distinct from the database glyph two rows up.
2026-09-22 00:30:56 +01:00
Alex 2f6b2ee8dd Fix a crash on the filtered agent routes, and two animation defects
AgentSection computed its folder breadcrumb in a useMemo placed below two
early returns for the empty states. A hook after an early return is
skipped on the render that takes it, which React rejects outright with
"rendered fewer hooks than expected" — the page showed only the error
boundary. Reachable now that each filter is its own route: landing
straight on an empty one renders once while the data loads and again
once it arrives empty, taking the early return the second time. The
memo moves above them.

Two problems in the chat's entrance animations, found while looking for
the reported flicker:

- .fade-in-bubble carried `opacity: 0` on the element and reached full
  opacity only by running `fadeInUp` to completion with `forwards`. The
  answer text was therefore visible *because* an animation had finished,
  so anything that stopped one running left it blank — including the
  obvious reduced-motion reset, which is presumably why the rule covers
  only .shimmer-text today. The start state moves into the keyframes,
  the element rests visible, and both entrances now honour
  prefers-reduced-motion.

- The timings were long for their jobs. An expanded tool call reaches
  its full height at once, so fading its content over half a second read
  as the content lagging the layout rather than as a reveal; it is now
  0.16s. The answer entrance goes to 0.26s with a 6px rise instead of
  0.5s and 10px.
2026-09-22 00:24:52 +01:00
Alex ac87430715 Animate the sidebar as a stack, and let it move on the click
The sidebar cross-faded between two states, which stopped describing
what was happening once sections could nest: entering a section slid,
but opening an agent from the agent list swapped in place with no
motion at all, so going deeper and going sideways looked identical.

Panels are now positioned from a single number — their depth relative
to the level on screen. A panel above the current level waits off to
the right, the current one sits at rest, and ones below park just off
to the left, so push and pop fall out of the same rule and no direction
has to be tracked. The panel behind travels a quarter of the width and
dims rather than sliding out with the one in front, and the arriving
panel carries a shadow off its leading edge that the container clips
once it lands, so the two read as stacked rather than adjacent.

The motion was also starting far too late. Mounting a section's page
costs a single ~170ms blocking frame in a production build, and the
sidebar's own class change rode along in that same commit: measured
from the click, the panels did not begin moving for ~290ms, so the
animation played to an audience that had stopped expecting it. The two
updates are now split by priority. The level lands as an urgent update
touching nothing but the sidebar, so React can commit and paint it
straight away; the route change goes through startTransition, which
renders the page at low priority and yields instead of blocking that
paint. The target's section is resolved from the path up front, so the
incoming panel arrives with its content already in place. The style
change now lands ~53ms after the click. Only translate and opacity are
animated, so the compositor keeps the motion smooth across the frames
the page render still costs.

Timing is tuned against where the travel actually lands rather than by
feel: half the distance by ~65ms so the panel tracks the click, 90% by
~180ms so the movement reads as movement, settled by ~300ms.
2026-09-21 23:45:50 +01:00
Alex cfbf61f4f3 Split the agents URL space and move its navigation into the sidebar
Routes under /agents covered two different things: using an agent (a
conversation) and managing them (the list, editor, logs, schedules).
Sharing the prefix left no way to tell them apart from the pathname, so
the sidebar could not react to one without also reacting to the other.

Management now lives under /agents/manage, and every caller builds its
links from agents/paths.ts rather than from a literal. Pre-split URLs
redirect, keeping their query.

With the prefixes distinct, agent management becomes a section like
settings and admin:

- The list's five filters are routes rather than component state, so a
  filtered view is linkable and survives a reload. The hand-rolled pill
  row is gone from the content on desktop.
- An agent's own pages (overview, logs, schedules) get a nav titled
  after the agent, replacing the breadcrumb-and-underline sub-nav.
  Sections nest to support it: `parentPath` makes back mean "up one
  level", so leaving an agent lands on the agent list rather than the
  chat.
- Sections named after a record are built per route, since the title
  comes from the store rather than the path.
- `pageTitle` distinguishes sections whose destinations are separate
  pages from ones whose destinations are views of a single page; the
  latter keep their own heading instead of flipping to "All".

Below lg, where the sidebar is an overlay, those view-style
destinations appear as a pill row on the page — bouncing out to an
index page to change a filter would be worse than a row of pills.

The workflow builder keeps its full-screen canvas and its own header,
the one place a content-owned nav still earns its keep.

Page shells converge on the shared padding, max width and header, and
the agent list's folder trail uses the shared breadcrumb primitives.
2026-09-21 23:04:12 +01:00
Alex 60532ec46c Move settings and admin navigation into the sidebar
Settings and admin both drove their pages from a horizontal tab strip.
Seven tabs no longer fit: settings had grown scroll arrows, gradient
masks and a hiddenGradient state machine just to survive on mobile, and
the active tab was resolved by comparing translated labels against the
URL. Teams and admin had no home in the strip at all — admin was
reachable only from the Help popover.

Replace both strips with a vertical nav that takes over the sidebar
while you are inside a section, plus a back button that returns to the
app. A declarative registry in navigation/sections.ts holds each
section's destinations; the active item is resolved from the route by
longest path match, so a detail route like /settings/tools/slack keeps
Tools active. Section state is derived from the route rather than
stored, so deep links and browser back keep working.

- Below lg the sidebar is an overlay, so /settings and /admin render
  their destination list as page content and each page carries a back
  link to it.
- Entering settings no longer clears the conversation; the back button
  returns to the route you came from and the chat list stays mounted,
  keeping its scroll position.
- Collapsing the sidebar inside a section shows the same destinations
  as icons instead of stranding you on one page.
- Detail views nested in a section page (a tool's config, a team) now
  use a breadcrumb rather than a second back arrow, so only the section
  nav means "leave".
- Teams and admin join the settings nav; admin-only entries are hidden
  from non-admins along with the group heading they leave empty.
2026-09-21 22:31:47 +01:00
Alex 295969b6a4 Merge pull request #2818 from arc53/branding-upd
logo upd
2026-09-21 18:00:18 +01:00
Alex a25bec6821 Merge remote-tracking branch 'origin/main' into branding-upd
# Conflicts:
#	docsgpt/core/models/anthropic.yaml
#	docsgpt/core/models/openai.yaml
#	frontend/src/admin/AdminUI.tsx
2026-09-21 17:52:53 +01:00
Alex 4755fee7c5 Merge pull request #2817 from arc53-machine/chore/zed-editor-config
chore: Zed project config, .editorconfig and shared pyright settings
2026-09-21 17:06:24 +01:00
Alex 9b85755c3b Merge pull request #2816 from arc53/feat/admin-quotas
feat: admin-set usage quotas per user and per team
2026-09-21 16:07:26 +01:00
Alex 1e14605ee7 fix(quotas): keyless agent chat bucket, keep disabled policies disabled, cached rates
- check_usage treats a request through a keyless (draft) agent as agent
  traffic, matching how its usage rows are bucketed and the headless rule
- dashboard edits carry the stored enabled flag instead of re-enabling the
  policy; disabled policies are labelled in the Quotas tab and the editor
- quota 429s send x-should-retry: false so OpenAI SDK clients do not retry
  a refusal that cannot succeed before the reset
- cached-input and cache-write rates for Anthropic, OpenRouter and Groq
  gpt-oss-120b; refresh OpenRouter deepseek-v3.2 list prices
- UsageQuota reuses usagePercent; docs note that a user override needs an
  existing user
2026-09-21 15:26:29 +01:00
Alex b5296df8a9 feat(pricing): cached-input rates for gpt-5.4-mini and gpt-5.4-nano
Checked against OpenAI's pricing page: gpt-5.5 at $5 / $30 (cached $0.50)
was already right. The mini and nano models declared no cached rate, so
cached prompt tokens were billed at the full input rate.
2026-09-21 14:30:48 +01:00
Alex 3a30f5cce9 feat(pricing): rates for the default DocsGPT model
$0.15 input, $0.50 output and $0.03 cached input per 1M tokens, so cost
budgets see usage of the default model instead of recording it at $0.
2026-09-21 13:36:22 +01:00
Alex 4bd259fc09 fix(quotas): address review: resume claims, agent bucket rule, UI races
- A tool continuation refused for usage now releases the resume claim it
  took; before, retries got a 409 until the stale claim was reverted.
- Agent traffic is any row with an agent key or an agent id, so keyless
  agents and workflow nodes count toward the agent bucket, not direct.
- The user quota modal discards responses for a previously opened user.
- The usage meter shows every limited bucket, not only 'all'.
- Restore the class separator on the analytics stat card that a formatter
  run removed, and align the OpenRouter DeepSeek description with its rates.
2026-09-21 12:44:26 +01:00
Alex 69f55b74cb refactor(quotas): validate policy bodies without exception text in responses
Validation problems are returned as values rather than raised and echoed
with str(exc), and a huge integer limit is rejected as out of range instead
of overflowing. Tests no longer call mutating endpoints inside asserts.
2026-09-21 12:15:42 +01:00
Alex 6db9014dfb fix(quotas): list unpriced models by recorded cost; integer token limits in status
The unpriced-model notice asked the live registry whether a model has a
price, so a priced model whose provider was later disabled showed up as
unpriced. It now lists models whose calls this period were all recorded at
$0. Token limits are serialized as integers.
2026-09-21 12:14:14 +01:00
Alex 1eacfdd3d0 docs: usage quotas
How the instance default, team allowances and user overrides resolve
(including users in several teams), the quota window, who is charged for
agent traffic, how cost budgets price models and what happens to unpriced
ones, and the admin and user API.
2026-09-21 12:11:38 +01:00
Alex 1c9c94eba7 feat(frontend): admin quota management, a usage meter and the quota chat error
The admin dashboard gets a Quotas tab for the instance default, team
allowances and user overrides, with a notice listing models that cost
limits cannot see. A Quota action on the Users tab shows a user's effective
limits, the layer each comes from and their usage, next to the editor for
their override. Each budget is either not set at that layer, a limit, or
unlimited.

Users with a quota see a usage meter with the reset time on the Analytics
settings page, and a refused chat request shows the used amount, the limit
and the reset time in the user's language.
2026-09-21 12:11:12 +01:00
Alex 21f43b2966 feat(quotas): admin quota API and GET /api/user/quota
Admins read and set the instance default, team allowances and user
overrides under /api/admin/quotas. A user's endpoint also returns the
limits those layers resolve to, the layer each came from and the usage
against them. The overview lists catalog models used this period that have
no price, since a cost limit cannot see them. Every write is audited.

GET /api/user/quota gives a user their own limited buckets, usage and reset
time without naming the policies behind them; any valid token may call it.
2026-09-21 12:11:12 +01:00
Alex 5406550bca feat(quotas): enforce user quotas on chat, agent, scheduled and webhook runs
check_usage now checks the billable user's quota on every request, before
the per-agent 24h limits, which keep applying to traffic through an agent.
Until now a request without an agent key skipped every limit. A refusal is
a 429 with Retry-After and a body naming the budget, usage, limit, the
layer the limit came from and when it resets.

Headless runs check the agent owner's quota before starting. A refused
scheduled run is recorded as budget_exceeded; a refused webhook run returns
a quota_exceeded result instead of raising, so Celery does not retry it.
2026-09-21 12:11:11 +01:00
Alex 6826313b60 feat(quotas): limit resolution and the quota service
docsgpt/quotas resolves a user's effective token and cost limits from the
policy rows that apply to them: their own override, then the most generous
allowance among their teams, then the instance default, then any defaults a
registered provider supplies. Each budget resolves on its own, and a team
counts once however many memberships the user holds in it.

QuotaService compares those limits with the user's token_usage totals over
the current QUOTA_PERIOD window (calendar-aligned, UTC, computed at read
time). It fails open, and skips the usage query for unlimited users.
2026-09-21 12:11:11 +01:00
Alex 6c42139224 fix(usage): stop double counting scheduled runs; attribute workflow node usage
sum_tokens_in_range now skips the scheduler's per-run rollup rows, whose
tokens are already on the run's per-call rows, so the per-agent 24h limit
and the admin total no longer count scheduled spend twice.

Workflow node LLMs carry the workflow agent's id, so their usage rows are
attributed to the agent instead of landing with a user id only.

usage_totals returns a user's tokens and cost since a window start, split
by interactive and agent-key traffic.
2026-09-21 12:11:11 +01:00
Alex 43fad2a865 feat(quotas): quota_policies table and a per-call cost on token_usage
Migration 0033 adds quota_policies (instance default, team per-member
allowance, user override; a token budget and a USD budget per row) and
token_usage.cost. The column is added IF NOT EXISTS so a database that
already carries it upgrades cleanly.

Every usage row now records the call's USD cost from the model catalog;
bring-your-own models are recorded at $0.
2026-09-21 12:11:11 +01:00
Alex b77561288d feat(pricing): per-million model rates and a cost module
Rename the unused *_cost_per_token capability fields to USD per 1M tokens,
add prompt-cache read/write rates, and ship list prices for the hosted
catalogs. The old per-token keys still load, scaled, with a warning.

docsgpt/pricing.py turns a call's token bins into a USD cost. Models with
no declared rate cost $0 unless QUOTA_UNPRICED_RATE_PER_MILLION is set.
2026-09-21 11:38:22 +01:00
Alex 1c1bc2538f Merge pull request #2812 from arc53-machine/feat/personal-access-tokens
feat: personal access tokens (scoped API tokens for CLI and CI/CD)
2026-09-21 10:55:33 +01:00
Alex 9ad09039e4 Merge pull request #2804 from arc53/fix/graphrag-extraction-and-retrieval
fix(graphrag): make graph retrieval and extraction work in a default install
2026-09-20 11:20:54 +01:00
Alex 63d93fc4b9 test(graphrag): pin that identical pages collapse into one
Review asked whether two document rows with the same text and metadata should
stay two pages. They should not: the caller gets four pages to hand a model,
and a crawl that ingested the same text twice would spend two of them on it.
The rows differ only by an id the model never sees. Pinned either way now.
2026-09-20 10:54:55 +01:00
Alex 692edcdd4e test(agents): give the graph tool tests the vector store they need
The tool now asks graphrag_available(), which wants pgvector as well as the
flag. One case set only the flag and passed locally off a dev .env, then
failed on CI's faiss default. An autouse fixture sets it for the module, so a
case that forgets fails for its own reason; the two about a different vector
store override it themselves.
2026-09-20 10:54:55 +01:00
Alex ef5b718f37 fix(graphrag): drop entities whose name normalizes to nothing
``canonical_name`` answers "" for a punctuation-only name, which callers are
meant to read as "no entity" -- ``_resolve_endpoint`` already does. Entity
extraction did not, and nodes merge on that key, so every such entity in a
source collapsed onto one shared node that belonged to none of them.
2026-09-20 10:45:05 +01:00
Alex 2b6d4d509e fix(graphrag): return a chunk once from entity_pages
The subject flag was in the GROUP BY, so a chunk two matching entities link --
one the page is about, one merely mentioned in it -- came back as two
identical pages and spent the caller's page budget twice on the same text.
It is aggregated with bool_or now, which is what the ordering wanted anyway.

Covered by a live test against a pgvector-shaped documents table: the graph
tables alone cannot answer this query, so nothing exercised it before.
2026-09-20 10:45:04 +01:00
Alex ccd8eb612f fix(agents): hand the graph tool's connection back, and gate it like the rest
Three faults in the graph tool, all on the agent's path:

The store was cached on the tool, and the executor caches the tool for the
whole agent run -- so one pgvector pooled connection stayed checked out across
every LLM round trip of that run, minutes at a time, and enough concurrent
runs exhaust the pool. GraphRAGRetriever releases its store before falling
back for this reason. The tool now releases it at the end of each action.

It gated on GRAPHRAG_ENABLED where everything else asks graphrag_available(),
which also requires the pgvector store. Under any other vector store the graph
tables are not the ones the sources were ingested into, but the tool was still
offered and still queried Postgres.

Pages were labelled by hand rather than through labels_from_metadata, which
exists so citation labels match across retrievers. A page read by the tool and
the same chunk retrieved by internal_search are one document, and citations
key on (source, title) -- so the research agent gave that document two
citation numbers. The recorded doc also keeps the full chunk text now, so it
dedupes against the retriever's copy; only what the model reads is truncated.
2026-09-20 10:45:04 +01:00
Alex bc0ef9f3b0 fix(retriever): search a graph source classically when its graph answers nothing
Only a raise routed a source to the ClassicRAG fallback, and every graph read
logs its own failure and returns empty. So a query that broke, a half-built
graph and a walk that genuinely found nothing were indistinguishable, and each
made the source contribute nothing at all to the answer -- no fallback, no
vector blend, which is skipped by the same early return.

A graph source that produces no documents now joins the classic batch, exactly
as a source with no graph already does.

The passage stage also rescanned every node's chunk list once per candidate.
It inverts the links once instead: 7.3 ms to 0.16 ms on a 400-node subgraph at
the candidate cap, with identical output.
2026-09-20 10:45:03 +01:00
Alex 8c5a5190d9 fix(retriever): carry a graph source's own options to the retriever
The three per-source graph options are read from the retrieval config the
Dispatcher hands over, and it only hands one over for a source it considers
overridden -- which it decided from chunks, score_threshold, rephrase_query
and prescreen alone. A source that changed only its graph options was not
"overridden", so nothing was carried and every graph source ran the defaults:
the UI toggles did nothing at all.

They count as an override now, for graphrag sources only. They mean nothing to
any other retriever, and an override also hands the source its own chunk
budget, which a classic source must not pick up from a graph setting.
2026-09-20 10:45:03 +01:00
Alex ecf02d0d60 fix(graphrag): take a source's write lock per chunk, and keep zero weights
Two builds of one source can overlap: the extraction lease is keyed by the
source's updated_at, and enabling a graph updates the source before it
dispatches, so a rebuild started while the last build runs gets a new key and
a lease of its own. Both builds could then pass a chunk's "done" check before
either committed and apply it twice -- doc_freq bumped twice, reproduced with
two live writers. A reset could also land in the middle of a chunk.

apply_chunk and delete_by_source now take a transaction-scoped advisory lock
keyed by the source before touching a row, as the schema bootstrap already
does for DDL. A single build's writes were already serial, so it loses
nothing; overlapping builds take turns chunk by chunk, and the second sees the
first's "done" row and returns (0, 0).

apply_chunk also still defaulted with `rel.get("weight") or 1.0`, turning an
explicit zero into a full-strength edge -- the conversion 3f774d81 removed from
add_edge and the ranker but missed here. Only a missing weight defaults now.
2026-09-19 20:27:27 +01:00
Alex b3d7e63ae3 fix(graphrag): compose graph SQL through psycopg instead of f-strings
Four graph-store queries formatted table and column names into the SQL
string: get_chunk_texts and delete_by_source (Bandit B608, alerts #582/#583
on main) and entity_pages/chunk_similarities from this branch (#661/#662,
dismissed). The names were validated by _safe_identifier, so none was
injectable, but each query was still a string built at runtime.

They are now fixed statements composed with psycopg.sql: identifiers go in as
sql.Identifier, values stay bound. Identifiers are lower-cased before quoting,
because PGVectorStore writes the same names unquoted and Postgres folds those
to lower case -- a quoted mixed-case name would address a different table.
Bandit reports nothing for docsgpt/graphrag now; the queries return the same
rows against a real graph as before.
2026-09-19 20:10:54 +01:00
Alex 558f803fed test(worker): send the real startup signals, and import celery_init one way
The lifecycle tests imported docsgpt.celery_init as a module next to the
file's from-imports, which code scanning flags. Patch the flag by path and send
the signals from celery.signals -- the objects a worker actually fires.
2026-09-19 14:53:18 +01:00
Alex 5e06470f9e test(graphrag): cover the graph reads, graph tool and vector blend without pgvector
CI has no pgvector, so every live graph test skips there and the queries this
branch added ran in no CI job at all. These pin what holds without a
database: the four new read queries bind every value (the entity name comes
from an LLM tool call) and map their rows; empty input runs no query; a failed
query returns nothing and releases its connection. The graph tool's plumbing,
the sources_have_graph gate and the hybrid path's vector ranking get the same.

Also drops a redundant chained comparison flagged by code scanning.
2026-09-19 14:33:14 +01:00
Alex b32d27c912 fix(agents): cite the pages the graph tool read
The graph tool records every page read_entity_pages returns in
retrieved_docs, but agents only ever collected internal_search's. A turn
answered from graph pages therefore emitted no sources: on the multi-hop e2e
run, a question answered from two read_entity_pages calls came back with none.

_search_tool_docs gathers both tools' documents the way the executor caches
them, and both the classic/agentic collector and the research agent's
per-step citations use it. Checked against the real executor and a real
graph: a graph-only turn now cites the pages it read.
2026-09-19 14:33:14 +01:00
Alex 877609dfb6 fix(worker): record worker state at startup, for every pool
in_worker() read task_join_will_block, which eventlet and gevent leave unset,
and the task's current_worker_task, which they scope to one greenlet -- so a
greenlet a task spawned in those pools still took the web-process branch and
dispatched to its own worker, where it could queue behind its parent and time
out.

The worker's own startup now records it: worker_init fires in every worker's
main process before the pool starts (where solo, threads, eventlet and gevent
run tasks, and what prefork children fork from), and worker_process_init in
each prefork child. worker_ready would be too late -- prefork children are
forked before it fires. The two existing checks stay for anything that runs
tasks without that startup.
2026-09-19 14:33:14 +01:00
Alex a20402a6dc fix(graphrag): give each extraction thread its own LLM
Provider-reported usage is kept on the LLM instance (_last_usage) and claimed
by whichever call finishes next. With GRAPHRAG_EXTRACTION_WORKERS > 1 (the
default is 8) every extraction thread shared one instance, so a call could
claim another call's provider counts while its own fell back to the estimate:
token_usage rows, and the cost they bill, could be attributed to the wrong
call and summed wrong.

Each pool thread now builds its own extraction LLM on first use. The calling
thread's instance is still built up front, so a misconfigured model fails the
run before any chunk is touched.
2026-09-19 14:33:14 +01:00
Alex 022cf69b49 fix(graphrag): fold an entity's singular and plural onto one key
canonical_name dropped "es" from every -ches/-ses/-zes plural, so "caches"
became "cach" while "cache" stayed "cache" -- the singular and plural landed on
two nodes, which is the split the function exists to prevent. The same rule
split the words documentation uses most: databases/database, responses/
response, releases/release, sizes/size.

An "-es" plural cannot say whether it is cache + "s" or batch + "es", so
instead of guessing, both sides now meet at the stem: the singular endings
-che/-she/-se/-ze/-xe drop their "e" the way the plurals drop "es", and a
singular's -ie folds to -y as -ies already did (cookie/cookies). The key is a
merge key that is never shown, so it only has to agree, not be a word. alias,
canvas, atlas and bias join the words that only look plural.

Found by the naming tests CI was missing: the module had no direct tests.
2026-09-19 14:33:14 +01:00
Alex 5285bb4115 fix(graphrag): give a failed chunk one more attempt before giving up
A chunk whose extraction failed was marked failed and the build moved on. The
checkpoint treats failed chunks as pending, but nothing ever reran the build,
so one transient error -- a provider hiccup, a single response that did not
parse -- left a permanent hole in the graph until someone rebuilt the whole
source.

Failed chunks now get one more attempt after the rest of the build, so a burst
of rate limiting has time to pass. Extraction errors, unparseable responses and
failed writes are all retried; a chunk is marked failed only when its retry
fails too, so it costs at most two calls. The chunks given up on are logged by
id, and progress still ends at the total.
2026-09-19 14:07:42 +01:00
Alex 15bdda8554 fix(worker): know you are in a worker from any thread, not only the task's
Celery records the executing task on the thread that runs it, so a thread that
task starts sees none. The embeddings client and read_document both decided
"am I in a worker?" from that alone, and from any other thread took the
web-process branch: dispatch to the worker they were running in and block on
the result. Celery refuses that get() ("Never call result.get() within a
task!"), so the embed failed and latched the 30s dispatch cooldown for every
caller after it; with joins allowed, read_document would instead wait on a
parsing queue only its own busy process serves.

Threads inside tasks are not hypothetical: per-source retrieval fans out to a
pool, so a scheduled or webhook agent searching several sources embedded from
pool threads. Graph extraction did too, which failed every chunk of a build.

in_worker() in celery_init answers for the whole process. Celery's
task_join_will_block is process-wide and set for every blocking pool (prefork,
solo, threads) -- exactly the condition under which dispatch-and-wait goes
wrong; eventlet/gevent leave it unset, so the task's own thread still counts
through current_worker_task. Verified with real workers on each blocking pool:
from a thread a task started, the old check dispatched and hit the error, the
new one embedded locally.
2026-09-19 14:07:42 +01:00
Alex 5e67f8c927 feat(settings): graph retrieval options in a graph source's retrieval settings
Exposes the three per-source graph options in the source's retrieval settings,
shown only when the retriever is graphrag: where the walk starts (entities or
relationships), whether passages join the walk, and whether vector hits are
blended in. Defaults match the backend's measured-best configuration and are
filled in for sources saved before the options existed. A note points at the
"search tool" exposure, which is what offers the graph to an agent.
2026-09-19 14:07:41 +01:00
Alex 6b9b193337 feat(agents): let agents walk a source's graph when they can search it
A one-shot graph ranking diffuses over the whole neighbourhood; a question
whose answer sits two hops away is better served by following the edges. The
graph_search tool gives an agent search_entities, get_relationships and
read_entity_pages over its graph sources. On a multi-hop corpus where the
bridging entity is never named, answers went from 0/10 with classic vector
retrieval to 10/10 with the tool, end to end through /stream.

The tool has no setting of its own. It is offered exactly where the agent can
already search: agentic and research agents always, a classic agent only for
sources exposed as a search tool. A graph source left at prefetch is used for
ranking only.

Tests now pin GRAPHRAG_ENABLED to its shipped default, as CI has: with a dev
.env enabling it, every agent test's graph check read the developer's real
database and left a pool to it behind.
2026-09-19 14:07:41 +01:00
Alex a83e1dc0af feat(graphrag): seed the walk from what entities are, and rank with passages and vector hits
Graph retrieval tied plain vector search at best and never beat it. Measured
across five corpora, the bottleneck was seeding, not the graph: the walk
started from nodes whose embeddings were computed from bare entity names, and
a whole question shares almost nothing with a name like "Quill".

Extraction now embeds each node from "name (type): description" and each
relationship as the fact it asserts ("Alder streams_to Quill: ..."), stored on
a new nullable graph_edges.fact_embedding column that ensure_vector_schema adds
in place. Entity names are canonicalised (case, punctuation, word breaks and a
cautious plural) so "VECTOR_STORE" and "vector stores" land on one node. Extraction calls run
concurrently (GRAPHRAG_EXTRACTION_WORKERS, default 8) while embedding and graph
writes stay serial on the task thread, so ordering and idempotency are
unchanged; that measured 8.4x faster with identical output.

Retrieval gains per-source options, stored under retrieval.graph and read live
at query time:

- seed_strategy: start from matching entities (default) or matching
  relationships, which can reach an entity the question never names;
- passage_nodes (on): walk the source's passages alongside entities, with
  PageRank damping 0.5 instead of 0.85;
- blend_vector (on): fuse the graph ranking with the source's vector ranking
  by reciprocal rank.

The defaults are the measured-best configuration. Through GraphRAGRetriever,
the new seeding moved recall@4 from 0.41 to 0.68 on a multi-hop corpus and
from 0.50 to 1.00 on the docs corpus, and regressed none of the corpora
measured. Existing graphs keep name-only embeddings until rebuilt.
2026-09-19 14:07:41 +01:00
Alex 0952a05aab Merge pull request #2805 from arc53/widget-license-upd 2026-09-17 22:17:59 +01:00
Alex b7e7872bf7 fix(tasks): a deferred duplicate records no result, rather than success
Returning a deferred marker traded one wrong signal for a worse one. A
redelivery reuses the original task id — Context.as_execution_options carries
task_id into the retry — so the duplicate's return marked the very id the
client polls as SUCCESS. /api/task_status reports celery's state verbatim and
the UI maps SUCCESS to "done", so the GraphRAG enable modal would announce a
finished build, rendered from a payload with no counts, while the run holding
the lease was still extracting.

Raise Ignore instead: celery records no state for the duplicate, so the task
id keeps whatever the holder sets and the poller keeps waiting. The autoretry
wrapper re-raises Ignore ahead of autoretry_for, so the wider
autoretry_for=(Exception,) on these tasks cannot turn it back into a retry.
2026-09-17 16:51:06 +01:00
Alex 3f774d813c fix(graphrag): review fixes — replay-safe writes, strict count, zero weights
Three findings from review, all on code this branch introduced.

apply_chunk was not replay-safe. commit() can report connection loss *after*
Postgres committed, and the reconnect retry then replays the write: _upsert_node
bumps doc_freq a second time and _add_edge inserts another row, since
graph_edges has no uniqueness constraint for a logical edge. The chunk's
graph_ingest_progress row is now written in the same transaction as the rows it
describes, and a replay that finds it already "done" returns (0, 0) without
touching the graph. Extraction drops its separate mark_chunk("done"): the
checkpoint and the graph can no longer disagree.

count_nodes swallows every query failure and answers 0, so extraction's
"fall back to the write count" handler could never run — a failed count after a
successful build reported an empty graph. count_nodes grows a strict mode that
re-raises; retrieval keeps the swallow, which is what routes a source to
ClassicRAG.

A zero edge weight was read as a full-strength link: `or 1.0` rewrote an
explicit 0 before the <= 0 filter. Only missing and null weights default now.
The same coercion sat in _ppr_scores, where it would have kept the ranker's
rule unreachable from the product path, so it is fixed there too.
2026-09-17 16:31:19 +01:00
Alex 9e9f130ed0 refactor(graphrag): spell the write retry out, and log a capped PPR run
Review feedback on _write_with_reconnect: an explicit return inside the loop
plus an implicit fall-through past it reads as a path that returns None. The
retry is exactly two attempts, so say so — first attempt, reconnect on
connection loss, second and final attempt — and every path now returns or
raises.

Also narrows the transitions hint to the (node, weight) tuples it holds, and
logs at debug when the power iteration hits its iteration cap instead of
returning the last iterate with no trace.
2026-09-17 16:03:39 +01:00
Alex 6ea737e57b fix(tasks): a deferred duplicate stands down instead of failing the task
A task that runs longer than the broker's visibility timeout is redelivered
while its first run is still going. The idempotency lease correctly stops the
duplicate from doing the work, but the duplicate then re-queued itself once
per LEASE_TTL until celery ran out of retries and raised
MaxRetriesExceededError — so a perfectly healthy long task (a large graph
extraction is the one that found this) reported a task failure, with a
traceback, while the real run was still making progress next to it.

Catch the exhaustion and return a "deferred" result instead. A normal
deferral still re-queues: only the give-up path changes, and the lease
holder's dedup row is left untouched so its own completion still records.
2026-09-17 15:49:56 +01:00
Alex e3d819d9fd fix(graphrag): stop losing chunks silently during a graph build
Three ways a chunk disappeared from a graph with no way to tell:

A build checks out one connection and then spends minutes per chunk waiting
on the model, so the connection idles long enough for the server or a pooler
to drop it. The pool only validates a connection when it hands one out, and
this one was handed out at the start of the build, so the next write raised
"the connection is lost", the chunk was marked failed, and the build carried
on a chunk short. apply_chunk and mark_chunk now reconnect and retry once;
every statement they run is an idempotent upsert, so a replay cannot
double-write. Only connection loss retries — a bad statement still surfaces.

An unparseable model response marked the chunk failed and logged nothing at
all, so failed_chunks was the only evidence and it named no chunk. Both
failure modes now log the chunk id.

The summary's node count summed per-chunk upserts, so an entity appearing in
ten chunks counted ten times: it reported writes, not graph size. It now
reports the distinct node count, falling back to the write count only if the
count query fails.
2026-09-17 15:44:23 +01:00
Alex 92e19ac177 fix(graphrag): dispatch extraction to the provider that serves the model
_build_extraction_llm passed settings.LLM_PROVIDER with the resolved
extraction model id. Those two disagree in any deployment that leaves the
provider at its default: the model id comes from GRAPHRAG_EXTRACTION_MODEL
or LLM_NAME, while the provider stays "docsgpt" — the hosted public
endpoint, which does not serve it. The request is rejected, the shared
fallback answers instead, and the graph gets built by a different model
than the one configured, with nothing in the summary saying so.

Resolve the provider from the model registry (owner-scoped, so per-user
BYOM ids resolve too), fall back to settings.LLM_PROVIDER only when the
model is unknown, and take the API key for the provider actually dispatched
to rather than the generic settings.API_KEY. The effective provider is
logged, since a silent swap was the whole failure mode.
2026-09-17 15:40:20 +01:00
Alex 94a33c4781 fix(graphrag): rank without scipy so graph retrieval stops falling back
networkx.pagerank delegates to a scipy implementation, and scipy is not a
DocsGPT dependency — it only arrives transitively through the optional
docling extra. In a default install every graph retrieval raised
ModuleNotFoundError inside _ppr_scores, hit the per-source except, and
degraded to ClassicRAG: the graph was built and paid for, then never used,
with one ERROR line per source per query as the only signal.

Rank with a local power iteration over the same row-normalized transition
matrix: undirected edges normalized per endpoint, dangling nodes
redistributed along the restart vector, and the restart vector normalized
across the nodes the subgraph actually holds so seed mass cannot leak.

Parity with networkx is asserted while scipy happens to be installed in the
test env, and the retrieval path is exercised with the import blocked.
2026-09-17 15:38:54 +01:00
Alex 24bcbd938e Merge pull request #2795 from arc53/ui-scan
UI refinements
2026-09-17 15:16:04 +01:00
Alex bc333f157c Merge pull request #2803 from arc53/feat/install-6-dev
feat: docsgpt dev, doctor and restart for the development loop
2026-09-17 14:48:25 +01:00
Alex c59741f604 Merge pull request #2800 from arc53/feat/install-5-native
feat: docsgpt up --native
2026-09-17 14:48:03 +01:00
Alex 9ca1eb8f34 Merge pull request #2799 from arc53/feat/install-4-backup
docsgpt backup and restore
2026-09-17 14:47:40 +01:00
Alex a8f2040eec Merge pull request #2802 from arc53-machine/refactor/settings-package
refactor(settings): split into per-domain modules, generate the settings reference
2026-09-17 13:39:02 +01:00
Alex 8b064c96f2 fix: keep credentials and 5000-digit paths out of the Redis URL errors
_redis_urls raises before any check runs and cli.main prints what it raises, so
every one of its five messages interpolated the URL it was given — password and
all — straight to stderr. They render it through _endpoint now.

_endpoint itself then echoed the whole path, so an over-long database number
came back at full length: 5132 characters of error for one bad setting. It caps
the path it renders, which protects every caller, since that string is written
into terminals, CI logs and error messages rather than reused as a URL.
2026-09-17 13:34:59 +01:00
Alex 6f981a5598 test: compare against the sanitiser, not a host substring
CodeQL flags "host" in message as incomplete URL sanitisation. I removed that
shape from two assertions and introduced a third in the same commit; this takes
it out of the provider test and the remaining Redis one, comparing with what
_endpoint produced instead, which is the contract those tests actually mean.
2026-09-17 13:19:34 +01:00
Alex da58c072a0 fix: five from the full review
- doctor printed OPENAI_BASE_URL raw, the third place a credential-bearing URL
  reached the terminal; it goes through _endpoint like the others.
- dev.run left its output readers unjoined, so a child's last lines could be
  lost on exit. The threads are kept and joined during teardown.
- logs read each file and then reopened it to follow, so anything written in
  between appeared in neither. One handle now serves both.
- dev took --port straight from argparse: 0 would have served on an ephemeral
  port while printing 0, and oversized values reach socket.bind. It goes
  through _port_number first.
- doctor --redis-url overrode only the broker, so a stale result backend or
  cache was still pinged and the flag looked broken. All three endpoints now
  come from the URL given.
2026-09-17 13:15:56 +01:00
Alex d0a0b352e0 fix: drop the host-substring assertions and explain the empty except
CodeQL flags `"host:port" in message` as incomplete URL sanitisation. It is a
message rather than a URL being authorised, so it is not a vulnerability, but
the check was red and comparing against the sanitiser's own output asserts the
real contract. The signal handler now says why it swallows: the child is
already gone, and shutdown must not fail on what it is cleaning up.
2026-09-17 13:03:59 +01:00
Alex 48285f4405 fix: doctor looks for alembic_version where alembic puts it
env.py sets no version_table_schema, so the table follows search_path. Pinning
public. in doctor's queries made them agree with each other and disagree with
alembic: on an install using another schema it reported no schema at all and
sent the user to migrate an already-migrated database. Both queries resolve the
table the same way alembic does now.
2026-09-17 12:51:50 +01:00
Alex bb50d0a91d fix: refuse to give the API the mock LLM's port
Both preflights passed for `dev --mock-llm --port 8090`: the port was free, and
then the mock and the API were each handed it. One child could not bind, and the
API was pointed at that port as its model server while trying to listen on it.
2026-09-17 12:39:41 +01:00
Alex 9e86dd4e94 fix: refuse a busy mock LLM port before dev starts anything
The mock starts first and the API and worker are pointed at it, but only the
API port was preflighted. A busy 8090 therefore showed up as a child exiting
once the rest were running, or as the API talking to whatever else was on that
port. --ui is deliberately left alone: vite.config.ts sets no strictPort, so
Vite moves to the next free port rather than failing.
2026-09-17 12:22:45 +01:00
Alex 66fbb11049 fix: doctor survives a Redis URL whose port will not parse
urlsplit accepts redis://host:notaport/0; only parts.port raises, and it raises
on access rather than at split time, so the ValueError fell outside the try.
_check_redis catches the client error and then formats it through _endpoint, so
doctor ended with a traceback from inside its own error path.
2026-09-17 12:12:28 +01:00
Alex 1cce6cfd67 test: cover the two things dev refuses to do
Running outside a checkout and starting on a port something else holds are both
refusals a developer will meet, and neither had a test. The second also asserts
that nothing is spawned when the port is taken, which is the part that matters:
the guard runs before any child process exists.
2026-09-17 12:02:40 +01:00
Alex 18313c7823 fix: keep credentials out of the text doctor borrows from its clients
Sanitising the URL in the message was not enough: the client's own error text
went into the detail too, and both psycopg and redis-py quote the URL they were
given. The reason is kept, the endpoint is kept, and the URL, username and
password are taken out of it.

The Redis test raised a generic error, so it asserted the password was absent
without ever exercising the path that leaked. Both tests now raise what the
clients actually raise.
2026-09-17 12:01:28 +01:00
Alex 9326c6edb1 fix: keep credentials out of doctor's output, and read the revision from the schema it checked
The Redis check put the whole URL in its failure message. A managed Redis URL
carries user:password@host, so a failed ping printed the password to the
terminal and into any log or issue the output was pasted into. It names
scheme://host:port/db now, and falls back to naming no URL when the value
cannot be parsed at all.

The Postgres check asked to_regclass about public.alembic_version and then read
version_num through search_path, so another schema could answer with a
different revision, or the query could fail, and doctor would send you to run
migrations against a database that is already fine.
2026-09-17 11:52:28 +01:00
Alex bcf2707efa fix: doctor reports a bad DOCSGPT_PORT instead of raising on it
_port_number was written so a hand-edited .env could not reach int() raw, and
then doctor did exactly that: a nonnumeric port ended the command with a
traceback rather than the message, in the one command whose job is to explain a
broken setup.
2026-09-17 11:43:53 +01:00