A turn whose agent raised wrote no user_logs row, so it only surfaced as
the agent's system error row. Every finished turn now writes its chat row,
at level error with the error when it failed, and linked to its trace; the
system row for the same traced activity is no longer listed twice.
DeepSeek and every other OpenAI-compatible API run through OpenAILLM, so
spans and metrics reported them as openai. The provider now comes from the
endpoint: known API hosts map to their gen_ai.provider.name (deepseek,
azure.ai.openai, mistral_ai, ...), and an OpenAI client pointed anywhere
else is openai_compatible. LLM spans also record server.address.
Paused, denied, skipped and client-run tool calls all go through
trace_unexecuted_tool_call. The research synthesis, continuation and
workflow step spans use with-blocks instead of hand-rolled GeneratorExit
handling. Fixes blank lines and adds a missing hint and docstrings.
Only message-linked traces cascaded, so scheduled runs, stateless /v1
rounds and unreserved turns kept their previews after the conversation was
deleted; they are now deleted by conversation_id. The Logs search and graph
branches filter by source, so the listing indexes lead with it.
The OTel replay and the trace INSERT ran in the stream's finally, so a slow
database held the SSE connection open after the last event. The trace is
still frozen when the stream ends, but written on a small writer pool.
Span.fail stored the exception message in span.error, which is stored and
exported whatever the content settings; a provider's content-filter error
can quote the prompt. span.error now holds the exception type, and the
message is a capture-gated preview, like tool errors already were. A
yielded stream error on the agent span is treated the same way.
With TRACES_ENABLED off, LLM calls record no GenAI metrics, and a streamed
answer is joined into a preview only when a span keeps it. Agent runs and
continuations build their invoke_agent span in one place, which now reads
the model from model_id. The trace panel shows a stream's model time next
to how long it was open. Adds missing type hints.
A failed trace-summary lookup now leaves the Logs page intact without
chips. Tool-call counts include only calls that ran, not their paused,
denied or skipped records. A local guardrail that fires unchanged on every
streamed segment is recorded once, so it cannot use up the span cap.
build_agent no longer takes request_id from the request body: it becomes
the primary LLM's usage request id, and quotas count distinct request ids,
so a client could make every call count as one. Requests refused after
setup started (unauthorized, over quota, resume conflict, setup error) now
write their trace, marked error, through an after-request hook; streaming
routes hand the trace to complete_stream instead.
The drawer is opened from a Logs row rather than a SheetTrigger, so Radix
had nothing to restore focus to. It now remembers what had focus when it
opened and returns there on close; the e2e spec checks it.
The drawer opens straight onto the trace. Its title stays as the dialog's
screen-reader name, and the unused subtitle string is removed from every
locale.
Colours come from theme tokens only: status pills and chips use muted,
primary and destructive; waterfall bars use chart-1..5, one per group of
span kinds, with destructive for failures and muted-foreground for
cancelled or pending steps. Rows select through ghost Buttons, the error box
is a destructive Alert and View trace is an outline Button. Bar geometry and
indentation move to CSS custom properties, scale ticks to static
positions, and the sheet takes layout classes only.
Rows open with Enter or Space and expose aria-expanded, so the View trace
action is reachable without a mouse. Durations round to whole seconds before
splitting, so 119.6 s reads 2m 00s instead of 1m 60s.
Tool results, denial comments and tool exception text now reach a span only
as a capture-gated preview; span.error is a fixed message, since it is
stored and exported regardless of content settings. A turn whose stream
yields an error event is recorded as failed. Search traces are listed by a
query copied into the small summary column, so the Logs timeline never
reads the spans JSONB. Adds GraphRAG span tests.
A tool-calling chat turn and a RAG turn are traced, linked from their Logs
rows and opened in the trace panel. The timeline scale starts at 0 and its
labels no longer wrap.
The trace lifecycle moves into a decorator so complete_stream keeps its
real parameters for introspection. GET /api/traces is readable with the
analytics:read scope, like the Logs endpoint it complements.
Covers what is recorded, where it is shown, the TRACES_* settings, the
gen_ai.* spans and metrics, content capture, and the limits of exporting
spans when the request ends.
Expanded Logs rows show the trace's duration, LLM calls, tokens, tool calls
and retrieval time, and open a side panel with a waterfall of every step:
retrieval, embeddings, per-source searches, model calls and tool calls,
with each step's details, previews and raw attributes. Chat turns paused
for approval show one waterfall per round. Searches and graph builds get
their own log types. Strings are translated in every locale.
GET /api/traces returns a Logs row's stored traces, scoped to the caller or
to an agent they own. get_user_logs rows now carry the ids that find their
trace and a merged summary (duration, LLM and tool calls, tokens, retrieval
time) from one batched lookup per page; searches, MCP searches and graph
builds, which have no log row of their own, are listed from their traces.
run_agent_headless records each unattended run under its endpoint; the
scheduler passes its run id and the webhook worker its task id so Logs rows
can find their trace, while the LLM's own request id stays untouched for
quota counts. /api/search and MCP search_docs record their retrieval, and a
graph build records every extraction call under one step.
StreamProcessor starts the trace and mints the request id before the
agent is built, so pre-fetch retrieval and compression are inside it and
side-channel LLM calls share the id. complete_stream activates it in the
SSE pump thread, binds the message and conversation, and writes it once
however the stream ends: paused, failed, abandoned or superseded (dropped).
user_logs rows now carry request_id and message_id.
Guardrail evaluations are traced when they call a remote check or fire; a
firing guardrail drops every content preview from the stored trace. Each
workflow node and each research phase is a step span, and the workflow run
id links the trace to its Logs row.
The dispatcher and each retriever's search open retrieval spans recording
sources, chunk counts, top scores and a chunk preview. The shared query
embedding, every per-source vector search and the prescreen rerank get
their own spans; pool workers carry the trace so they nest correctly.
@log_activity opens the invoke_agent span for every agent run and names
the trace's activity; gen_continuation opens its own. ToolExecutor.execute
records each executed call with redacted argument and result previews, and
paused, denied, skipped and client-executed calls are recorded where they
are decided.
The token-usage wrappers open one span per decorated invocation, so a
primary attempt, its retry and a fallback appear as siblings with the
provider that ran. Stream spans start on the first pull, carry tokens,
cost, time to first token and cache hits, and feed the GenAI metrics.
Finished traces are replayed with their recorded timestamps and explicit
parents under the request's server span, using gen_ai.* attribute names.
Content is exported only when OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT
opts in. gen_ai.client token-usage and operation-duration histograms are
recorded per model call.
docsgpt.tracing records agent, LLM, tool, retrieval and embedding spans in
memory and writes the finished trace once. Container spans nest by a
per-thread stack so suspended generators cannot corrupt the tree; open
spans are cancelled at flush and previews are bounded and redacted.
One row per execution holding its span tree as JSONB, linked to Logs rows
by request, message, activity and workflow-run ids. message_id cascades so
deleting or truncating a conversation drops its traces; a daily beat task
enforces TRACES_RETENTION_DAYS.
TRACES_* settings for the per-request trace timeline and its GenAI OTel
export. opentelemetry-api was only transitive; the tracing package
imports it directly.
`float("inf") > 0` is true, so a stored `"inf"` was preserved as if it
were a real count and reached the UI as "∞". Require a finite positive
value, which also covers `nan` explicitly rather than by accident.
Both chunk headers inlined the same `token_count ? … : '-'` expression, and
`ChunkType` declared every metadata value a string even though the JSON-typed
stores return the count as a number.
Move the formatting into `formatChunkTokens`, which accepts either shape and
only falls back to a dash when the value is genuinely unusable, and type
`ChunkType.metadata` to match what the API actually returns.
`reembed_wiki_page_worker` built each chunk's metadata from scratch as
`{source, title, filename}`, discarding everything the chunker attached
to the document -- including the per-chunk `token_count` every other
ingest path records. Wiki chunks therefore reached the source viewer
with no count at all.
Carry `extra_info` over and let the page's own path and title win on top,
so a chunk that inherited a stale source from an earlier conversion is
still keyed and filtered correctly.
Chunk cards in the source viewer read `metadata.token_count` and print a
bare "-" whenever it is absent. Ingestion records the count, but chunks
indexed before it was written -- and any path that rebuilds a chunk's
metadata from scratch -- reach the UI without it, so the whole file shows
"-" with no way to recover the number short of a re-ingest.
Fill the key in on the way out when it is missing or unusable (empty,
non-numeric, zero or negative), counting only the page being returned so
the cost is bounded by `per_page`. A count that is already recorded is
left untouched, including a numeric string, since stores round-trip
metadata types differently.
The fallback counts in cl100k rather than the embedding model's
tokenizer: loading that tokenizer can reach for a Hugging Face download,
which does not belong in a request path.
The DEFAULT 'unknown' added in the last pass stopped a previous-release
process from raising NotNullViolation mid-rollout, but it threw the
attribution away to do it. The highest-volume auth_events writers are OIDC
login and silent renewal, where the actor is simply the user, so a rollout
window would have flattened exactly the rows that were trivially recoverable.
Lifts both derivation rules into SQL functions -- auth_events_derive_actor
and auth_events_derive_target -- so the backfill and a BEFORE INSERT trigger
share one definition rather than two copies of the same CASE drifting apart.
The trigger guards on NEW.actor_id IS NULL, which identifies a legacy insert
exactly because the repository always supplies one. That distinction matters
for target_id: a current writer sets it NULL deliberately for events with no
user target, and deriving there would name the actor as their own target.
Kept rather than scheduled for removal: it is a no-op on the path the
repository takes, and it keeps the column's contract true for any writer that
bypasses it.
Correctness
- /api/remote never recorded source.created, so URL, GitHub and connector
sources had a source.deleted with no matching creation. All three creation
paths now go through one _audit_source_created helper.
- The prompt-cache rate divided cached tokens by a whole bucket's prompt
tokens. A bucket is a day and mixes calls whose provider reports a cache
breakdown with calls whose provider does not, so filtering buckets in the
client could not separate them and the rate was understated by however much
traffic ran on a non-reporting provider. The denominator is now computed in
SQL over the reporting rows.
- The outcome pill matched values nothing writes. Guardrails emit triggered /
not_evaluated and the device feed emits dispatched; the map had blocked /
denied / allowed, so a guardrail that fired rendered neutral grey -- the one
signal the merged feed exists to surface. Fixtures were seeding the
fictional values, so the tests passed on it too.
- Stream duration_ms timed the consumer. stream_token_usage is a generator,
so start-to-exhaustion includes the agent loop's tool handling and the SSE
client's pace; a slow browser recorded ~30s for a sub-second call. It now
accumulates only the time spent inside next().
Safety
- Activity filters failed open: an unknown facet or unparseable timestamp was
dropped, and no filter means every row, so a typo widened an audit view and
on the export streamed the full history. Both are now a 400.
- The search term was interpolated into an ILIKE pattern, so "100%" matched
everything and "q1_report" matched more than it should. Escaped.
- 0034 set actor_id NOT NULL with no default. A previous-release process
inserting mid-rollout would raise, and in admin/routes.py that insert shares
the request transaction, so a role grant beside it would roll back too.
Noise and dead code
- The per-user panel is a security panel: data-plane events file under the
actor, so an active account's routine deletes pushed a denied login out of
the 20-row window. It now excludes them; the Activity tab shows everything.
- device_audit_log had no created_at-leading index, so the merged feed
sequentially scanned that branch every page (migration 0036).
- conversation.deleted_all no longer records when nothing was deleted, and
agent.updated no longer records an empty field list.
- Dropped by_model from /admin/usage (no consumer; an extra aggregate per page
load), the duplicate filter surface on AuthEventsRepository that nothing
called, and the unreachable FLOW_LABELS.schedule entry.
- Type hints on record_event's conn and the remaining unannotated helpers.
- 0034's backfill set target_id to the acting admin for instance- and
team-scoped quota policy changes, which are filed under the actor and have
no user target. It now returns NULL for those, matching what quotas.py
records going forward, with a test covering all three scopes.
- The CSV export wrote every cell verbatim. user_agent is an attacker's raw
header, recorded without authenticating on a denied login, and the export
is opened in a spreadsheet by an admin -- a cell starting "=" would be
evaluated there. Every cell now has a leading formula trigger neutralized.
- The activity feed applied every response it received, so a slow reply for
an old filter could overwrite the current one. Guarded by a request id,
the same way settings/Analytics already does.
- The activity detail expander and the top-user drill-down were row onClick
handlers, unreachable without a mouse. Both are real buttons now, the
expander carrying aria-expanded and a label naming its event.
- FLOW_LABELS was applied to both breakdowns in the spend modal, so a model
named "fallback" or "workflow" rendered as a flow description. Only the
flow table maps keys now.
- Changing the range in that modal left the previous range's totals and
chart on screen while the new request was in flight.
The comment claimed X-Export-Max-Rows says whether the cap was reached; it
reports the cap that was in effect, which is what lets a caller tell a
truncated export from a complete one.
SCIM provisioning is performed by the identity provider on a user, so the
rows now carry actor_id='system:scim' and the user as target_id. The
assertions pin both rather than ignoring the new columns.
_page_after decided whether to append " AND " or " WHERE " by looking for
"WHERE" in the SQL it had just been handed, which only worked because no
union branch contains one. The query builder now always emits a WHERE, so
adding a branch with its own predicate cannot silently break the cursor.
bucketed_totals now returns cost and cached_tokens, so the exact-equality
assertion in test_day_bucket_sums_per_day had to grow the two fields; it
pins that an unreported cache bin reads NULL rather than 0.
The per-user breakdown test was seeded with a source='schedule' row, which
is a run-level rollup and no longer counted. Reseeded on a real flow, and
both spend queries gained a test that a rollup is excluded.
Review pass over the preceding commits.
- The activity export paged by OFFSET. The journals are append-only and the
feed is newest-first, so rows written mid-export shift the window down and
repeat rows already emitted. Pages by keyset on (created_at, feed, id) now.
- top_token_users and the per-user breakdown counted run-level rollup rows.
A scheduled run already has a row per LLM call, so its spend was billed
twice -- invisible while the column was tokens, obvious once it was dollars.
Both now exclude them, matching every other spend query.
- total_cost was summed from the per-model split, which drops rows with no
model_id and so undercounted the figure printed above the series. Summed
from the series instead.
- The CSV export encoded its detail cell without the fallback the NDJSON
branch had, so a non-JSON-native value would have failed the stream.
- The activity view reset pagination in an effect, which fetched the stale
page against the new filters before fetching again. Reset in the setters.
- The audit taxonomy is no longer re-exported from the helper module; the one
caller that wanted it imports from where it lives.