The DEFAULT 'unknown' added in the last pass stopped a previous-release
process from raising NotNullViolation mid-rollout, but it threw the
attribution away to do it. The highest-volume auth_events writers are OIDC
login and silent renewal, where the actor is simply the user, so a rollout
window would have flattened exactly the rows that were trivially recoverable.
Lifts both derivation rules into SQL functions -- auth_events_derive_actor
and auth_events_derive_target -- so the backfill and a BEFORE INSERT trigger
share one definition rather than two copies of the same CASE drifting apart.
The trigger guards on NEW.actor_id IS NULL, which identifies a legacy insert
exactly because the repository always supplies one. That distinction matters
for target_id: a current writer sets it NULL deliberately for events with no
user target, and deriving there would name the actor as their own target.
Kept rather than scheduled for removal: it is a no-op on the path the
repository takes, and it keeps the column's contract true for any writer that
bypasses it.
Correctness
- /api/remote never recorded source.created, so URL, GitHub and connector
sources had a source.deleted with no matching creation. All three creation
paths now go through one _audit_source_created helper.
- The prompt-cache rate divided cached tokens by a whole bucket's prompt
tokens. A bucket is a day and mixes calls whose provider reports a cache
breakdown with calls whose provider does not, so filtering buckets in the
client could not separate them and the rate was understated by however much
traffic ran on a non-reporting provider. The denominator is now computed in
SQL over the reporting rows.
- The outcome pill matched values nothing writes. Guardrails emit triggered /
not_evaluated and the device feed emits dispatched; the map had blocked /
denied / allowed, so a guardrail that fired rendered neutral grey -- the one
signal the merged feed exists to surface. Fixtures were seeding the
fictional values, so the tests passed on it too.
- Stream duration_ms timed the consumer. stream_token_usage is a generator,
so start-to-exhaustion includes the agent loop's tool handling and the SSE
client's pace; a slow browser recorded ~30s for a sub-second call. It now
accumulates only the time spent inside next().
Safety
- Activity filters failed open: an unknown facet or unparseable timestamp was
dropped, and no filter means every row, so a typo widened an audit view and
on the export streamed the full history. Both are now a 400.
- The search term was interpolated into an ILIKE pattern, so "100%" matched
everything and "q1_report" matched more than it should. Escaped.
- 0034 set actor_id NOT NULL with no default. A previous-release process
inserting mid-rollout would raise, and in admin/routes.py that insert shares
the request transaction, so a role grant beside it would roll back too.
Noise and dead code
- The per-user panel is a security panel: data-plane events file under the
actor, so an active account's routine deletes pushed a denied login out of
the 20-row window. It now excludes them; the Activity tab shows everything.
- device_audit_log had no created_at-leading index, so the merged feed
sequentially scanned that branch every page (migration 0036).
- conversation.deleted_all no longer records when nothing was deleted, and
agent.updated no longer records an empty field list.
- Dropped by_model from /admin/usage (no consumer; an extra aggregate per page
load), the duplicate filter surface on AuthEventsRepository that nothing
called, and the unreachable FLOW_LABELS.schedule entry.
- Type hints on record_event's conn and the remaining unannotated helpers.
- 0034's backfill set target_id to the acting admin for instance- and
team-scoped quota policy changes, which are filed under the actor and have
no user target. It now returns NULL for those, matching what quotas.py
records going forward, with a test covering all three scopes.
- The CSV export wrote every cell verbatim. user_agent is an attacker's raw
header, recorded without authenticating on a denied login, and the export
is opened in a spreadsheet by an admin -- a cell starting "=" would be
evaluated there. Every cell now has a leading formula trigger neutralized.
- The activity feed applied every response it received, so a slow reply for
an old filter could overwrite the current one. Guarded by a request id,
the same way settings/Analytics already does.
- The activity detail expander and the top-user drill-down were row onClick
handlers, unreachable without a mouse. Both are real buttons now, the
expander carrying aria-expanded and a label naming its event.
- FLOW_LABELS was applied to both breakdowns in the spend modal, so a model
named "fallback" or "workflow" rendered as a flow description. Only the
flow table maps keys now.
- Changing the range in that modal left the previous range's totals and
chart on screen while the new request was in flight.
The comment claimed X-Export-Max-Rows says whether the cap was reached; it
reports the cap that was in effect, which is what lets a caller tell a
truncated export from a complete one.
SCIM provisioning is performed by the identity provider on a user, so the
rows now carry actor_id='system:scim' and the user as target_id. The
assertions pin both rather than ignoring the new columns.
_page_after decided whether to append " AND " or " WHERE " by looking for
"WHERE" in the SQL it had just been handed, which only worked because no
union branch contains one. The query builder now always emits a WHERE, so
adding a branch with its own predicate cannot silently break the cursor.
bucketed_totals now returns cost and cached_tokens, so the exact-equality
assertion in test_day_bucket_sums_per_day had to grow the two fields; it
pins that an unreported cache bin reads NULL rather than 0.
The per-user breakdown test was seeded with a source='schedule' row, which
is a run-level rollup and no longer counted. Reseeded on a real flow, and
both spend queries gained a test that a rollup is excluded.
Review pass over the preceding commits.
- The activity export paged by OFFSET. The journals are append-only and the
feed is newest-first, so rows written mid-export shift the window down and
repeat rows already emitted. Pages by keyset on (created_at, feed, id) now.
- top_token_users and the per-user breakdown counted run-level rollup rows.
A scheduled run already has a row per LLM call, so its spend was billed
twice -- invisible while the column was tokens, obvious once it was dollars.
Both now exclude them, matching every other spend query.
- total_cost was summed from the per-model split, which drops rows with no
model_id and so undercounted the figure printed above the series. Summed
from the series instead.
- The CSV export encoded its detail cell without the fallback the NDJSON
branch had, so a non-JSON-native value would have failed the stream.
- The activity view reset pagination in an effect, which fetched the stale
page against the new filters before fetching again. Reset in the setters.
- The audit taxonomy is no longer re-exported from the helper module; the one
caller that wanted it imports from where it lives.
Covers what the previous commits changed: the actor_id / target_id split and
the system actors, the data-plane events, the merged Activity feed and its
export, and the new usage and per-user spend endpoints.
Also corrects "there is no UI for these events yet" in the OIDC login
auditing section, which is no longer true.
Three gaps in one view.
Cost: token_usage.cost has been written on every call since quotas landed, and
tokens_by_model already returned it, but bucketed_totals and top_token_users
selected tokens only. An admin could set a USD quota in /admin/quotas and had
no way to see the spend it was capping. Buckets and top users now carry cost,
the endpoint reports a window total and a per-model split, and the chart takes
a Tokens/Cost toggle with a currency axis.
Group-by: the endpoint has supported group_by=model|agent|source from the
start and the UI only ever sent bucket=day. The selector is now wired, and a
grouped series pads missing buckets so a model that was idle on Tuesday plots
a zero instead of shifting its whole row one bar left.
Latency: surfaced from the columns the previous commit added, as p50/p95 with
a median time-to-first-token underneath. Unmeasured rows are excluded rather
than counted as zero, and the card says so when nothing was measured.
Also surfaces the prompt-cache hit rate, computed only over rows whose
provider reported a cache breakdown -- NULL means "not reported", and folding
those in as 0% would understate it.
Per-user drill-down: GET /api/admin/users/<id>/usage returns a daily series
plus splits by model and by flow, reachable from the top-users table and from
the user detail dialog. The detail dialog previously showed one tokens_30d
number, which answers neither "what is this person costing" nor "what is
driving it".
The wrappers in docsgpt/usage.py already measured how long each call took --
for the llm_gen_finished / llm_stream_finished log lines -- and then threw the
number away. Nothing in the schema recorded it, so an operator could see what
an instance spent but never how slow it was.
Adds token_usage.duration_ms and token_usage.ttft_ms, and threads the
measurements the wrappers already take into the insert. The stream wrapper now
also stamps the moment of the first yielded chunk.
Both columns are nullable, and deliberately so: ttft_ms is NULL for a
non-streaming call and for a stream that failed before yielding anything, and
every row written before this migration is NULL too. A 0 there would drag a
p50 toward an instant first token that never happened.
Replaces the Audit tab with a merged Activity view over all three journals.
The old table rendered four columns -- event, user, IP, when -- and dropped
everything else the row carried. auth_events captures a metadata object and a
user agent on every insert, and neither was ever shown, so "Admin revoked ·
user-42" could not tell you which admin. Rows now expand into a detail panel
carrying the actor, the affected user, the outcome, the user agent and every
metadata field.
Filters were an event box whose placeholder was the only hint that denied
logins are spelled "oidc_login_denied", plus a user-id box. The backend
supported a time window that the UI never sent. Now: a category facet, an
event picker fed by the instance's own catalogue, quick time ranges and a
search that reaches into the detail payload. Export buttons pull the filtered
feed down as CSV or NDJSON.
The deep link from Overview (?event=oidc_login_denied) still lands pre-
filtered, and the route stays /admin/audit so existing bookmarks work.
Drops the two single-journal service readers the merged feed supersedes --
getDeviceAudit had shipped months ago and never had a caller. The
/api/admin/audit endpoint stays for API consumers.
TableCell/TableHeader gain colSpan so the detail row can span the table
instead of a second table being invented for it.
DocsGPT keeps three append-only journals -- auth_events, device_audit_log and
guardrail_events -- each readable only on its own terms, and the second of
those had an admin endpoint that no UI ever called. An operator reviewing an
incident wants one timeline, not three.
Adds GET /api/admin/activity: all three journals projected onto a common row
shape and merged, ordered newest first, with facets for journal, category,
event name, actor, affected user, a time window and free-text search over the
detail payload. A category filter that only one journal can satisfy drops the
others from the union rather than scanning and discarding them.
GET /api/admin/activity/events returns the distinct (event, category) pairs
the instance has actually recorded, so the filter UI can offer real choices
instead of asking an operator to type "oidc_login_denied" from memory.
GET /api/admin/activity/export streams the same filtered feed as CSV or
NDJSON. Streamed rather than buffered, and capped, so a compliance export of a
busy instance is bounded work.
guardrail_events.api_key (a raw agent key) and matched_value (unredacted
source text) are excluded from the projection, mirroring the exclusion the
per-agent guardrail view already makes. Admin-gating is not a reason to widen
what a list response carries.
Both fail on main as of f5b02ed3, before any of this branch's changes —
verified by running them against main in a clean worktree, where they fail
with byte-identical errors. Neither is testing a broken feature; both hang
their assertion on a DOM landmark that #2819 moved.
C13 waited for AgentPageHeader's `Agent sub-navigation` nav on
/agents/manage/logs/:agentId. That page now renders CurrentSectionHeader +
SectionPills, and WorkflowBuilder is the only remaining caller of
AgentPageHeader, so the landmark never appears. C2 proved a locale switch
had taken effect by looking for the Settings entry as a `link` role, which
60532ec4 moved into the sidebar as something else.
The behaviour underneath both still works — the logs page renders its
panels, and switching locale swaps the strings and persists to
localStorage. Only the selectors are stale, so these could be repaired
rather than dropped if someone wants the coverage back.
Removed insertStubAgent with C13, its only caller, and the two imports
that went with it.
Logins, role grants and provisioning were audited from the first release;
creating and deleting sources, agents, agent keys and conversations were not.
An operator reviewing the trail could see who signed in but not who deleted
the source they were asking about.
Adds docsgpt/api/audit.py: one helper that records an event inside the
caller's transaction, in a savepoint, swallowing failures -- an audit write
must never be able to fail the action it describes -- and tolerating the
absence of a Flask request so a Celery task can record too.
Events added: source.created (upload and wiki), source.deleted,
source.reingested, agent.created, agent.updated, agent.deleted,
agent.key_regenerated, conversation.deleted, conversation.deleted_all.
agent.updated records field names only, never values, which can carry prompts
and credentials.
The same module carries the event -> category map (identity / access / config
/ data) that the admin activity feed filters on.
auth_events.user_id was overloaded and meant something different per call
site: admin mutations filed the row under the target user and hid the acting
admin in metadata->>'by', team events filed it under the actor, and quota
events switched between the two depending on scope. The practical consequence
was that "show me everything admin X did" had no answer.
Adds actor_id (never NULL) and target_id (NULL when the event is not about a
user) alongside the existing column, which keeps its meaning as the per-user
feed key so nothing historical is rewritten. The backfill recovers the actor
from metadata for rows that predate the columns.
Also adds the indexes the cross-user admin feed needs. The only index was
(user_id, created_at DESC), so the global feed -- which orders by created_at
with no user predicate -- fell back to a sequential scan plus a sort on every
page.
The feed gains actor, multi-event, until and free-text search filters, plus a
distinct-event catalogue so the UI can offer what the instance recorded
instead of asking operators to type a name from memory.
test_each_provider_has_expected_ids asserts the exact set of model ids per
provider against a snapshot. For the first-party catalogs that is what we
want: the snapshot exists so an upstream id typo fails before it silently
breaks every agent that references the old id.
openai_compatible is different. It is the zero-Python extension point —
the README's "Add an OpenAI-compatible provider" section and
examples/mistral.yaml.example both tell you to drop a YAML next to the
built-ins, and the loader globs the whole directory. So anyone who follows
those instructions gets a permanently red test on their fork, for having
done the documented thing. The _internal suffix carve-out only covers the
one filename we happen to gitignore.
Treat the snapshot as a floor for that provider: the built-in ids must all
still be present, extra ones are somebody else's provider. The rename this
guards against still drops an id, so it still fails, and the assertion
names the missing ids instead of printing two large sets to diff by eye.
Verified by copying examples/mistral.yaml.example into the catalog
directory: red before, green after, and still red when a built-in id is
renamed.
Three things the UI pass started but did not carry all the way.
The MCP server dialog's Basic Authentication branch was fixed to use the
label keys for its labels and the placeholder keys for its placeholders,
but the API Key and Bearer Token branches above it in the same
renderAuthFields were left reading placeholders.apiKey and
placeholders.bearerToken as their <Label>. Both fields showed the
placeholder text twice over: the label above the input read "Your secret
API key" where the input already hinted it. authTypes.apiKey and
authTypes.bearer already held "API Key" and "Bearer Token". (Note that
the Basic branch that prompted all this is unreachable — `basic` is
commented out of the authTypes list, so that form cannot be opened.)
Dropped a dead `t(...) || 'Scopes (comma separated)'` fallback in the same
file while there. t() returns the key string when a key is missing, never
a falsy value, so the right-hand side could not run; the branch would
render the raw key rather than the English it looks like it guards.
The plural fix covered English only. resultSummary, view_more,
actionsFound and needsSetup all took _one/_other in en.json, but de, es
and ru kept a single form, so "1 Chunks abgerufen", "1 fragmentos
recuperados" and "ещё 1 источников" still rendered — the exact bug the
change set out to remove, in every language but the one it was tested in.
Added _one/_other for de and es, and _one/_few/_many/_other for ru, which
needs the three-way split (1 источник, 2 источника, 5 источников).
Translated needsSetup while there; it was still English in all three.
jp, zh and zh-TW have a single plural category, so their base form is
already correct and stays as it is.
Replacing Spinner with a skeleton dropped the role="status" that Spinner
carried, and the skeleton divs announce nothing, so a screen reader got
no signal at all between navigating to Agents and the list appearing.
Put it on SkeletonLoader itself rather than at the call sites, so every
variant gets it. It is sr-only, which is position: absolute, so it never
becomes a flex or grid item in the containers these render into.
The branch forked before the agents navigation moved into the sidebar, so
main had since rewritten both files this PR touches. Two conflicts.
AgentsList: main split the filters into their own routes and dropped the
FILTER_TABS pill row and the page <h1> for CurrentSectionHeader +
SectionPills, so the page-level skeleton this branch added no longer has
a place to live — it keyed off `activeFilter === 'all'` and gated the
whole page on all four sections having loaded. Resolved to main's
structure, which already loads each section independently, and kept only
the part that was the point: the section's Spinner becomes the agentCards
skeleton. That also drops `isLoading && !isDataLoaded`, which would have
suppressed the indicator entirely on a refetch over cached data.
Keeping per-section loading is the better behaviour anyway. The page-level
loader hid the section headers and the New Folder / Import / New Agent
buttons for as long as the slowest of the four endpoints took, and hid the
sections that had already arrived along with them.
AgentLogs: main moved the page onto CurrentSectionHeader + SectionPills
and a max-w-6xl wrapper. Took that, with this branch's removal of the two
page-level spinners — Analytics, GuardrailEvents and Logs each have their
own skeleton, so the outer spinners were the second of the two loaders the
PR set out to remove. Gated on `agentId` alone rather than
`agentId && (loadingAgent || agent)`: the children only ever needed the id,
and the longer condition mounted all three, then unmounted them mid-flight
when the agent fetch failed, leaving the page blank with no error.
The hook-order crash this branch still carried in AgentSection is fixed on
main already (the breadcrumb useMemo sits above the early returns), and the
merge takes that.
The previous commit claimed both handlers, but only the first one
changed: the search text expected `if (...) return;` on a single line
while prettier had already wrapped it across two, so the second
replacement silently matched nothing. A ctrl/cmd/shift-click on the
collapsed-list entry therefore still closed the mobile nav and cleared
the selected agent in the tab it left behind.
main has been red since 295969b: the model catalog gained claude-fable-5,
gpt-5.6-sol and the alibaba and zai providers, but the pieces that track
the catalog were not updated alongside them. Five tests across
test_model_registry_yaml.py and test_pricing.py have been failing since.
- EXPECTED_IDS picks up the two new hosted models and the two new
openai_compatible ones. The snapshot is deliberate — it exists to
catch an unintended rename — so it is updated, not relaxed.
- test_everything_set sets DASHSCOPE_API_KEY and ZAI_API_KEY alongside
the DEEPSEEK one it already set. Every openai_compatible catalog reads
its own key from the environment, so without them "everything set"
quietly excluded the two new providers.
- claude-fable-5 and gpt-5.6-sol had no rates, which
test_hosted_builtin_models_are_priced requires of every hosted model,
and which token_usage.cost and the quota ceilings both read. Priced on
the existing ladder: fable-5 above Opus 4.7 at 10/50 per million with
cached input at a tenth and cache writes at 1.25x, matching the other
Anthropic entries; sol alongside the gpt-5.5 flagship at 5/30.
- Opening a folder dropped the active filter. agentsListPath() always
built off the bare /agents/manage root, so from /agents/manage/mine it
navigated to /agents/manage?folder=..., and filterFromPath() then read
back `all`. A regression from making the filters routable: they used
to be component state, which a replace-navigation left alone. The
helper now takes the filter and builds off agentsFilterPath(), which
also fixes the other half — switching filter while inside a folder
re-navigates to the filtered path with the folder kept, instead of the
sync effect pulling it back to the unfiltered URL.
- Both manage-agents handlers cleared the selected agent and closed the
mobile nav before returning early for a modifier-click, so opening the
list in a new tab mutated the tab you stayed on. The modifier check
and preventDefault() now run first.
- The detail breadcrumbs had the name precedence inverted. Tools.tsx
shows `customName || displayName` throughout, so a renamed tool kept
its old name in the breadcrumb alone. ToolConfig now matches, and
RemoteDeviceConfig leads with device?.name as its own heading does.
Boxes is the densest glyph in the settings nav by some way — nine
elements and 43 path commands against a median of about twelve. Three
cubes with their internal facets inside 20px puts the strokes almost on
top of each other, so at the nav's 1.75 weight it renders as a darker
blob than the icons above and below it, even though the stroke width is
identical. Blocks sits at the median, matching Sources, Tools and Teams,
and stays distinct from the database glyph two rows up.
AgentSection computed its folder breadcrumb in a useMemo placed below two
early returns for the empty states. A hook after an early return is
skipped on the render that takes it, which React rejects outright with
"rendered fewer hooks than expected" — the page showed only the error
boundary. Reachable now that each filter is its own route: landing
straight on an empty one renders once while the data loads and again
once it arrives empty, taking the early return the second time. The
memo moves above them.
Two problems in the chat's entrance animations, found while looking for
the reported flicker:
- .fade-in-bubble carried `opacity: 0` on the element and reached full
opacity only by running `fadeInUp` to completion with `forwards`. The
answer text was therefore visible *because* an animation had finished,
so anything that stopped one running left it blank — including the
obvious reduced-motion reset, which is presumably why the rule covers
only .shimmer-text today. The start state moves into the keyframes,
the element rests visible, and both entrances now honour
prefers-reduced-motion.
- The timings were long for their jobs. An expanded tool call reaches
its full height at once, so fading its content over half a second read
as the content lagging the layout rather than as a reveal; it is now
0.16s. The answer entrance goes to 0.26s with a 6px rise instead of
0.5s and 10px.
The sidebar cross-faded between two states, which stopped describing
what was happening once sections could nest: entering a section slid,
but opening an agent from the agent list swapped in place with no
motion at all, so going deeper and going sideways looked identical.
Panels are now positioned from a single number — their depth relative
to the level on screen. A panel above the current level waits off to
the right, the current one sits at rest, and ones below park just off
to the left, so push and pop fall out of the same rule and no direction
has to be tracked. The panel behind travels a quarter of the width and
dims rather than sliding out with the one in front, and the arriving
panel carries a shadow off its leading edge that the container clips
once it lands, so the two read as stacked rather than adjacent.
The motion was also starting far too late. Mounting a section's page
costs a single ~170ms blocking frame in a production build, and the
sidebar's own class change rode along in that same commit: measured
from the click, the panels did not begin moving for ~290ms, so the
animation played to an audience that had stopped expecting it. The two
updates are now split by priority. The level lands as an urgent update
touching nothing but the sidebar, so React can commit and paint it
straight away; the route change goes through startTransition, which
renders the page at low priority and yields instead of blocking that
paint. The target's section is resolved from the path up front, so the
incoming panel arrives with its content already in place. The style
change now lands ~53ms after the click. Only translate and opacity are
animated, so the compositor keeps the motion smooth across the frames
the page render still costs.
Timing is tuned against where the travel actually lands rather than by
feel: half the distance by ~65ms so the panel tracks the click, 90% by
~180ms so the movement reads as movement, settled by ~300ms.
Routes under /agents covered two different things: using an agent (a
conversation) and managing them (the list, editor, logs, schedules).
Sharing the prefix left no way to tell them apart from the pathname, so
the sidebar could not react to one without also reacting to the other.
Management now lives under /agents/manage, and every caller builds its
links from agents/paths.ts rather than from a literal. Pre-split URLs
redirect, keeping their query.
With the prefixes distinct, agent management becomes a section like
settings and admin:
- The list's five filters are routes rather than component state, so a
filtered view is linkable and survives a reload. The hand-rolled pill
row is gone from the content on desktop.
- An agent's own pages (overview, logs, schedules) get a nav titled
after the agent, replacing the breadcrumb-and-underline sub-nav.
Sections nest to support it: `parentPath` makes back mean "up one
level", so leaving an agent lands on the agent list rather than the
chat.
- Sections named after a record are built per route, since the title
comes from the store rather than the path.
- `pageTitle` distinguishes sections whose destinations are separate
pages from ones whose destinations are views of a single page; the
latter keep their own heading instead of flipping to "All".
Below lg, where the sidebar is an overlay, those view-style
destinations appear as a pill row on the page — bouncing out to an
index page to change a filter would be worse than a row of pills.
The workflow builder keeps its full-screen canvas and its own header,
the one place a content-owned nav still earns its keep.
Page shells converge on the shared padding, max width and header, and
the agent list's folder trail uses the shared breadcrumb primitives.
Settings and admin both drove their pages from a horizontal tab strip.
Seven tabs no longer fit: settings had grown scroll arrows, gradient
masks and a hiddenGradient state machine just to survive on mobile, and
the active tab was resolved by comparing translated labels against the
URL. Teams and admin had no home in the strip at all — admin was
reachable only from the Help popover.
Replace both strips with a vertical nav that takes over the sidebar
while you are inside a section, plus a back button that returns to the
app. A declarative registry in navigation/sections.ts holds each
section's destinations; the active item is resolved from the route by
longest path match, so a detail route like /settings/tools/slack keeps
Tools active. Section state is derived from the route rather than
stored, so deep links and browser back keep working.
- Below lg the sidebar is an overlay, so /settings and /admin render
their destination list as page content and each page carries a back
link to it.
- Entering settings no longer clears the conversation; the back button
returns to the route you came from and the chat list stays mounted,
keeping its scroll position.
- Collapsing the sidebar inside a section shows the same destinations
as icons instead of stranding you on one page.
- Detail views nested in a section page (a tool's config, a team) now
use a breadcrumb rather than a second back arrow, so only the section
nav means "leave".
- Teams and admin join the settings nav; admin-only entries are hidden
from non-admins along with the group heading they leave empty.
- .zed/settings.json: ruff + basedpyright for Python (no format on save, the
tree is not ruff-format clean), ESLint fixes then Prettier for the frontend,
scan exclusions for caches and build outputs, .jwt_secret_key as private
- .zed/tasks.json: dev services, API, worker, frontend, pytest/vitest for the
current file or test, linting, uv lock + requirements export
- .zed/debug.json: debugpy targets matching .vscode/launch.json
- .editorconfig: whitespace rules for every editor
- [tool.pyright] in pyproject.toml: venv, import root and excludes shared by
pyright, basedpyright and Pylance
- .gitignore: track only the shared files under .zed/
- CONTRIBUTING: editor setup section; fix the stale ESLint config path
- check_usage treats a request through a keyless (draft) agent as agent
traffic, matching how its usage rows are bucketed and the headless rule
- dashboard edits carry the stored enabled flag instead of re-enabling the
policy; disabled policies are labelled in the Quotas tab and the editor
- quota 429s send x-should-retry: false so OpenAI SDK clients do not retry
a refusal that cannot succeed before the reset
- cached-input and cache-write rates for Anthropic, OpenRouter and Groq
gpt-oss-120b; refresh OpenRouter deepseek-v3.2 list prices
- UsageQuota reuses usagePercent; docs note that a user override needs an
existing user
Checked against OpenAI's pricing page: gpt-5.5 at $5 / $30 (cached $0.50)
was already right. The mini and nano models declared no cached rate, so
cached prompt tokens were billed at the full input rate.
- A tool continuation refused for usage now releases the resume claim it
took; before, retries got a 409 until the stale claim was reverted.
- Agent traffic is any row with an agent key or an agent id, so keyless
agents and workflow nodes count toward the agent bucket, not direct.
- The user quota modal discards responses for a previously opened user.
- The usage meter shows every limited bucket, not only 'all'.
- Restore the class separator on the analytics stat card that a formatter
run removed, and align the OpenRouter DeepSeek description with its rates.
Validation problems are returned as values rather than raised and echoed
with str(exc), and a huge integer limit is rejected as out of range instead
of overflowing. Tests no longer call mutating endpoints inside asserts.
The unpriced-model notice asked the live registry whether a model has a
price, so a priced model whose provider was later disabled showed up as
unpriced. It now lists models whose calls this period were all recorded at
$0. Token limits are serialized as integers.
How the instance default, team allowances and user overrides resolve
(including users in several teams), the quota window, who is charged for
agent traffic, how cost budgets price models and what happens to unpriced
ones, and the admin and user API.
The admin dashboard gets a Quotas tab for the instance default, team
allowances and user overrides, with a notice listing models that cost
limits cannot see. A Quota action on the Users tab shows a user's effective
limits, the layer each comes from and their usage, next to the editor for
their override. Each budget is either not set at that layer, a limit, or
unlimited.
Users with a quota see a usage meter with the reset time on the Analytics
settings page, and a refused chat request shows the used amount, the limit
and the reset time in the user's language.
Admins read and set the instance default, team allowances and user
overrides under /api/admin/quotas. A user's endpoint also returns the
limits those layers resolve to, the layer each came from and the usage
against them. The overview lists catalog models used this period that have
no price, since a cost limit cannot see them. Every write is audited.
GET /api/user/quota gives a user their own limited buckets, usage and reset
time without naming the policies behind them; any valid token may call it.
check_usage now checks the billable user's quota on every request, before
the per-agent 24h limits, which keep applying to traffic through an agent.
Until now a request without an agent key skipped every limit. A refusal is
a 429 with Retry-After and a body naming the budget, usage, limit, the
layer the limit came from and when it resets.
Headless runs check the agent owner's quota before starting. A refused
scheduled run is recorded as budget_exceeded; a refused webhook run returns
a quota_exceeded result instead of raising, so Celery does not retry it.