Commit Graph
5598 Commits
Author SHA1 Message Date
Alex 91d670b7bb Merge pull request #2825 from arc53-machine/feat/genai-traces
Execution traces: GenAI OpenTelemetry spans and a trace waterfall in Logs
2026-09-24 00:12:12 +01:00
arc53-machine f2d92bf8df Log failed chat turns as chat entries
A turn whose agent raised wrote no user_logs row, so it only surfaced as
the agent's system error row. Every finished turn now writes its chat row,
at level error with the error when it failed, and linked to its trace; the
system row for the same traced activity is no longer listed twice.
2026-09-24 00:08:53 +01:00
arc53-machine 3156ee37a1 Name the provider a model call actually reached
DeepSeek and every other OpenAI-compatible API run through OpenAILLM, so
spans and metrics reported them as openai. The provider now comes from the
endpoint: known API hosts map to their gen_ai.provider.name (deepseek,
azure.ai.openai, mistral_ai, ...), and an OpenAI client pointed anywhere
else is openai_compatible. LLM spans also record server.address.
2026-09-24 00:08:53 +01:00
arc53-machine 6b145d7aad Wait for the background trace write in the /v1 replay test
The first request's trace is written on the trace-writer pool after the
response returns; assert once it has landed rather than racing it.
2026-09-23 23:23:41 +01:00
arc53-machine a74420ed2c Trace unexecuted tool calls in one place; tidy span lifecycles
Paused, denied, skipped and client-run tool calls all go through
trace_unexecuted_tool_call. The research synthesis, continuation and
workflow step spans use with-blocks instead of hand-rolled GeneratorExit
handling. Fixes blank lines and adds a missing hint and docstrings.
2026-09-23 22:42:05 +01:00
arc53-machine b3aa509d3a Delete all of a conversation's traces and index trace listing
Only message-linked traces cascaded, so scheduled runs, stateless /v1
rounds and unreserved turns kept their previews after the conversation was
deleted; they are now deleted by conversation_id. The Logs search and graph
branches filter by source, so the listing indexes lead with it.
2026-09-23 22:42:05 +01:00
arc53-machine 248ebc8050 Discard the setup trace of a replayed /v1 request
An Idempotency-Key retry returns the cached response after setup already
ran; its trace was then written as a failure no Logs row points to.
2026-09-23 22:42:05 +01:00
arc53-machine 778830208b Write chat traces off the stream's thread
The OTel replay and the trace INSERT ran in the stream's finally, so a slow
database held the SSE connection open after the last event. The trace is
still frozen when the stream ends, but written on a small writer pool.
2026-09-23 22:42:05 +01:00
arc53-machine 0591580014 Keep exception text out of stored and exported span errors
Span.fail stored the exception message in span.error, which is stored and
exported whatever the content settings; a provider's content-filter error
can quote the prompt. span.error now holds the exception type, and the
message is a capture-gated preview, like tool errors already were. A
yielded stream error on the agent span is treated the same way.
2026-09-23 22:42:05 +01:00
arc53-machine 0bbbac0eb2 Skip trace work when disabled; share the agent span builder
With TRACES_ENABLED off, LLM calls record no GenAI metrics, and a streamed
answer is joined into a preview only when a span keeps it. Agent runs and
continuations build their invoke_agent span in one place, which now reads
the model from model_id. The trace panel shows a stream's model time next
to how long it was open. Adds missing type hints.
2026-09-23 22:09:58 +01:00
arc53-machine 699906a015 Give a scheduled run's trace to the user who scheduled it
On a shared agent the run's Logs row belongs to the scheduling user, not
the agent owner, so the trace is stored under that user.
2026-09-23 22:09:58 +01:00
arc53-machine 9f7f0b2f1e Keep trace summaries from breaking Logs and fix their counts
A failed trace-summary lookup now leaves the Logs page intact without
chips. Tool-call counts include only calls that ran, not their paused,
denied or skipped records. A local guardrail that fires unchanged on every
streamed segment is recorded once, so it cannot use up the span cap.
2026-09-23 22:09:58 +01:00
arc53-machine 24796ae61f Mint request ids on the server and trace refused requests
build_agent no longer takes request_id from the request body: it becomes
the primary LLM's usage request id, and quotas count distinct request ids,
so a client could make every call count as one. Requests refused after
setup started (unauthorized, over quota, resume conflict, setup error) now
write their trace, marked error, through an after-request hook; streaming
routes hand the trace to complete_stream instead.
2026-09-23 22:09:58 +01:00
arc53-machine 82b31760ee Put View trace before the trace chips in a Logs row 2026-09-23 21:08:27 +01:00
arc53-machine 5272440a21 Return focus to View trace when the drawer closes
The drawer is opened from a Logs row rather than a SheetTrigger, so Radix
had nothing to restore focus to. It now remembers what had focus when it
opened and returns there on close; the e2e spec checks it.
2026-09-23 21:08:27 +01:00
arc53-machine 56e8437a83 Drop the visible heading from the trace drawer
The drawer opens straight onto the trace. Its title stays as the dialog's
screen-reader name, and the unused subtitle string is removed from every
locale.
2026-09-23 21:06:49 +01:00
arc53-machine d7319928d1 Build the trace UI on theme tokens and existing components
Colours come from theme tokens only: status pills and chips use muted,
primary and destructive; waterfall bars use chart-1..5, one per group of
span kinds, with destructive for failures and muted-foreground for
cancelled or pending steps. Rows select through ghost Buttons, the error box
is a destructive Alert and View trace is an outline Button. Bar geometry and
indentation move to CSS custom properties, scale ticks to static
positions, and the sheet takes layout classes only.
2026-09-23 20:53:18 +01:00
arc53-machine 8157bcf3bf Make Logs rows keyboard-operable and fix minute rounding
Rows open with Enter or Space and expose aria-expanded, so the View trace
action is reachable without a mouse. Durations round to whole seconds before
splitting, so 119.6 s reads 2m 00s instead of 1m 60s.
2026-09-23 18:11:19 +01:00
arc53-machine efec1a2f83 Keep tool output out of span errors and fix trace status and listing
Tool results, denial comments and tool exception text now reach a span only
as a capture-gated preview; span.error is a fixed message, since it is
stored and exported regardless of content settings. A turn whose stream
yields an error event is recorded as failed. Search traces are listed by a
query copied into the small summary column, so the Logs timeline never
reads the spans JSONB. Adds GraphRAG span tests.
2026-09-23 18:11:19 +01:00
arc53-machine 285fdc914c Add an e2e spec for execution traces; tidy the timeline scale
A tool-calling chat turn and a RAG turn are traced, linked from their Logs
rows and opened in the trace panel. The timeline scale starts at 0 and its
labels no longer wrap.
2026-09-23 17:55:30 +01:00
arc53-machine 3512685728 Keep complete_stream's signature and classify /api/traces for tokens
The trace lifecycle moves into a decorator so complete_stream keeps its
real parameters for introspection. GET /api/traces is readable with the
analytics:read scope, like the Logs endpoint it complements.
2026-09-23 17:50:19 +01:00
arc53-machine 4e1002f5e0 Document execution traces and the GenAI OpenTelemetry export
Covers what is recorded, where it is shown, the TRACES_* settings, the
gen_ai.* spans and metrics, content capture, and the limits of exporting
spans when the request ends.
2026-09-23 17:46:18 +01:00
arc53-machine 3310b7fa1b Show execution traces in the Logs UI
Expanded Logs rows show the trace's duration, LLM calls, tokens, tool calls
and retrieval time, and open a side panel with a waterfall of every step:
retrieval, embeddings, per-source searches, model calls and tool calls,
with each step's details, previews and raw attributes. Chat turns paused
for approval show one waterfall per round. Searches and graph builds get
their own log types. Strings are translated in every locale.
2026-09-23 17:45:19 +01:00
arc53-machine 26881ba66b Serve traces to the Logs UI
GET /api/traces returns a Logs row's stored traces, scoped to the caller or
to an agent they own. get_user_logs rows now carry the ids that find their
trace and a merged summary (duration, LLM and tool calls, tokens, retrieval
time) from one batched lookup per page; searches, MCP searches and graph
builds, which have no log row of their own, are listed from their traces.
2026-09-23 17:39:09 +01:00
arc53-machine 5d0992eef8 Trace scheduled, webhook, search, MCP and graph-extraction runs
run_agent_headless records each unattended run under its endpoint; the
scheduler passes its run id and the webhook worker its task id so Logs rows
can find their trace, while the LLM's own request id stays untouched for
quota counts. /api/search and MCP search_docs record their retrieval, and a
graph build records every extraction call under one step.
2026-09-23 17:37:24 +01:00
arc53-machine bae842d151 Trace every chat turn from agent setup to the last event
StreamProcessor starts the trace and mints the request id before the
agent is built, so pre-fetch retrieval and compression are inside it and
side-channel LLM calls share the id. complete_stream activates it in the
SSE pump thread, binds the message and conversation, and writes it once
however the stream ends: paused, failed, abandoned or superseded (dropped).
user_logs rows now carry request_id and message_id.
2026-09-23 17:34:54 +01:00
arc53-machine da1bc16eb6 Trace guardrails, workflow steps and research phases
Guardrail evaluations are traced when they call a remote check or fire; a
firing guardrail drops every content preview from the stored trace. Each
workflow node and each research phase is a step span, and the workflow run
id links the trace to its Logs row.
2026-09-23 17:29:21 +01:00
arc53-machine f9605e7794 Trace RAG retrieval, query embeddings and prescreen
The dispatcher and each retriever's search open retrieval spans recording
sources, chunk counts, top scores and a chunk preview. The shared query
embedding, every per-source vector search and the prescreen rerank get
their own spans; pool workers carry the trace so they nest correctly.
2026-09-23 17:27:43 +01:00
arc53-machine a15940c281 Trace agent runs and tool calls
@log_activity opens the invoke_agent span for every agent run and names
the trace's activity; gen_continuation opens its own. ToolExecutor.execute
records each executed call with redacted argument and result previews, and
paused, denied, skipped and client-executed calls are recorded where they
are decided.
2026-09-23 17:25:41 +01:00
arc53-machine e8166e6c29 Record a chat span for every LLM call
The token-usage wrappers open one span per decorated invocation, so a
primary attempt, its retry and a fallback appear as siblings with the
provider that ran. Stream spans start on the first pull, carry tokens,
cost, time to first token and cache hits, and feed the GenAI metrics.
2026-09-23 17:23:06 +01:00
arc53-machine b5e9257659 Export traces as OpenTelemetry GenAI spans and metrics
Finished traces are replayed with their recorded timestamps and explicit
parents under the request's server span, using gen_ai.* attribute names.
Content is exported only when OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT
opts in. gen_ai.client token-usage and operation-duration histograms are
recorded per model call.
2026-09-23 17:21:25 +01:00
arc53-machine 3f45d019b0 Add the per-request trace recorder
docsgpt.tracing records agent, LLM, tool, retrieval and embedding spans in
memory and writes the finished trace once. Container spans nest by a
per-thread stack so suspended generators cannot corrupt the tree; open
spans are cancelled at flush and previews are bounded and redacted.
2026-09-23 17:19:54 +01:00
arc53-machine f3f817b1e8 Store execution traces in a new request_traces table
One row per execution holding its span tree as JSONB, linked to Logs rows
by request, message, activity and workflow-run ids. message_id cascades so
deleting or truncating a conversation drops its traces; a daily beat task
enforces TRACES_RETENTION_DAYS.
2026-09-23 17:19:54 +01:00
arc53-machine f6269c483f Add execution-trace settings and declare opentelemetry-api
TRACES_* settings for the per-request trace timeline and its GenAI OTel
export. opentelemetry-api was only transitive; the tracing package
imports it directly.
2026-09-23 17:14:51 +01:00
Alex 8a9f3a75d6 Merge pull request #2823 from arc53-machine/fix-chunk-token-counts
Show token counts on source chunks
2026-09-22 16:45:42 +01:00
arc53-machine db8db62c93 Treat a non-finite token count as unrecorded
`float("inf") > 0` is true, so a stored `"inf"` was preserved as if it
were a real count and reached the UI as "∞". Require a finite positive
value, which also covers `nan` explicitly rather than by accident.
2026-09-22 14:01:27 +01:00
arc53-machine 3954de9b9e Render chunk token counts through one tested helper
Both chunk headers inlined the same `token_count ? … : '-'` expression, and
`ChunkType` declared every metadata value a string even though the JSON-typed
stores return the count as a number.

Move the formatting into `formatChunkTokens`, which accepts either shape and
only falls back to a dash when the value is genuinely unusable, and type
`ChunkType.metadata` to match what the API actually returns.
2026-09-22 13:45:40 +01:00
arc53-machine d787f54c2d Keep the chunker's metadata when re-embedding a wiki page
`reembed_wiki_page_worker` built each chunk's metadata from scratch as
`{source, title, filename}`, discarding everything the chunker attached
to the document -- including the per-chunk `token_count` every other
ingest path records. Wiki chunks therefore reached the source viewer
with no count at all.

Carry `extra_info` over and let the page's own path and title win on top,
so a chunk that inherited a stale source from an earlier conversion is
still keyed and filtered correctly.
2026-09-22 13:44:24 +01:00
arc53-machine 0aecdc32c9 Backfill missing chunk token counts in /api/get_chunks
Chunk cards in the source viewer read `metadata.token_count` and print a
bare "-" whenever it is absent. Ingestion records the count, but chunks
indexed before it was written -- and any path that rebuilds a chunk's
metadata from scratch -- reach the UI without it, so the whole file shows
"-" with no way to recover the number short of a re-ingest.

Fill the key in on the way out when it is missing or unusable (empty,
non-numeric, zero or negative), counting only the page being returned so
the cost is bounded by `per_page`. A count that is already recorded is
left untouched, including a numeric string, since stores round-trip
metadata types differently.

The fallback counts in cl100k rather than the embedding model's
tokenizer: loading that tokenizer can reach for a Hugging Face download,
which does not belong in a request path.
2026-09-22 13:43:45 +01:00
Alex cf3ae3cdea Merge pull request #2822 from arc53-machine/feat/admin-activity-and-usage
feat(admin): activity feed, data-plane audit, and spend/latency in usage
2026-09-22 13:32:50 +01:00
arc53-machine f693111de1 fix(audit): derive attribution for legacy inserts instead of defaulting it
The DEFAULT 'unknown' added in the last pass stopped a previous-release
process from raising NotNullViolation mid-rollout, but it threw the
attribution away to do it. The highest-volume auth_events writers are OIDC
login and silent renewal, where the actor is simply the user, so a rollout
window would have flattened exactly the rows that were trivially recoverable.

Lifts both derivation rules into SQL functions -- auth_events_derive_actor
and auth_events_derive_target -- so the backfill and a BEFORE INSERT trigger
share one definition rather than two copies of the same CASE drifting apart.

The trigger guards on NEW.actor_id IS NULL, which identifies a legacy insert
exactly because the repository always supplies one. That distinction matters
for target_id: a current writer sets it NULL deliberately for events with no
user target, and deriving there would name the actor as their own target.

Kept rather than scheduled for removal: it is a no-op on the path the
repository takes, and it keeps the column's contract true for any writer that
bypasses it.
2026-09-22 13:22:11 +01:00
arc53-machine 25c82003d7 fix(admin): second review pass
Correctness
- /api/remote never recorded source.created, so URL, GitHub and connector
  sources had a source.deleted with no matching creation. All three creation
  paths now go through one _audit_source_created helper.
- The prompt-cache rate divided cached tokens by a whole bucket's prompt
  tokens. A bucket is a day and mixes calls whose provider reports a cache
  breakdown with calls whose provider does not, so filtering buckets in the
  client could not separate them and the rate was understated by however much
  traffic ran on a non-reporting provider. The denominator is now computed in
  SQL over the reporting rows.
- The outcome pill matched values nothing writes. Guardrails emit triggered /
  not_evaluated and the device feed emits dispatched; the map had blocked /
  denied / allowed, so a guardrail that fired rendered neutral grey -- the one
  signal the merged feed exists to surface. Fixtures were seeding the
  fictional values, so the tests passed on it too.
- Stream duration_ms timed the consumer. stream_token_usage is a generator,
  so start-to-exhaustion includes the agent loop's tool handling and the SSE
  client's pace; a slow browser recorded ~30s for a sub-second call. It now
  accumulates only the time spent inside next().

Safety
- Activity filters failed open: an unknown facet or unparseable timestamp was
  dropped, and no filter means every row, so a typo widened an audit view and
  on the export streamed the full history. Both are now a 400.
- The search term was interpolated into an ILIKE pattern, so "100%" matched
  everything and "q1_report" matched more than it should. Escaped.
- 0034 set actor_id NOT NULL with no default. A previous-release process
  inserting mid-rollout would raise, and in admin/routes.py that insert shares
  the request transaction, so a role grant beside it would roll back too.

Noise and dead code
- The per-user panel is a security panel: data-plane events file under the
  actor, so an active account's routine deletes pushed a denied login out of
  the 20-row window. It now excludes them; the Activity tab shows everything.
- device_audit_log had no created_at-leading index, so the merged feed
  sequentially scanned that branch every page (migration 0036).
- conversation.deleted_all no longer records when nothing was deleted, and
  agent.updated no longer records an empty field list.
- Dropped by_model from /admin/usage (no consumer; an extra aggregate per page
  load), the duplicate filter surface on AuthEventsRepository that nothing
  called, and the unreachable FLOW_LABELS.schedule entry.
- Type hints on record_event's conn and the remaining unannotated helpers.
2026-09-22 12:47:40 +01:00
Alex 6c5dfbe9a0 Merge pull request #2807 from simpleqt/sq919/e2e-plan-link
docs(e2e): drop the link to the never-committed e2e-plan.md
2026-09-22 11:38:47 +01:00
Alex 0c85d766dd Merge pull request #2820 from arc53/ui-scan
UI refinements
2026-09-22 11:22:43 +01:00
arc53-machine 2ea9d66789 fix(admin): address review findings
- 0034's backfill set target_id to the acting admin for instance- and
  team-scoped quota policy changes, which are filed under the actor and have
  no user target. It now returns NULL for those, matching what quotas.py
  records going forward, with a test covering all three scopes.
- The CSV export wrote every cell verbatim. user_agent is an attacker's raw
  header, recorded without authenticating on a denied login, and the export
  is opened in a spreadsheet by an admin -- a cell starting "=" would be
  evaluated there. Every cell now has a leading formula trigger neutralized.
- The activity feed applied every response it received, so a slow reply for
  an old filter could overwrite the current one. Guarded by a request id,
  the same way settings/Analytics already does.
- The activity detail expander and the top-user drill-down were row onClick
  handlers, unreachable without a mouse. Both are real buttons now, the
  expander carrying aria-expanded and a label naming its event.
- FLOW_LABELS was applied to both breakdowns in the spend modal, so a model
  named "fallback" or "workflow" rendered as a flow description. Only the
  flow table maps keys now.
- Changing the range in that modal left the previous range's totals and
  chart on screen while the new request was in flight.
2026-09-22 11:13:11 +01:00
arc53-machine 64200a7f23 docs(admin): say what the export cap header actually reports
The comment claimed X-Export-Max-Rows says whether the cap was reached; it
reports the cap that was in effect, which is what lets a caller tell a
truncated export from a complete one.
2026-09-22 10:52:55 +01:00
arc53-machine bddc5e67af test: assert the SCIM audit attribution
SCIM provisioning is performed by the identity provider on a user, so the
rows now carry actor_id='system:scim' and the user as target_id. The
assertions pin both rather than ignoring the new columns.
2026-09-22 10:46:54 +01:00
arc53-machine 350caf6ace refactor(admin): make the activity keyset clause unconditional
_page_after decided whether to append " AND " or " WHERE " by looking for
"WHERE" in the SQL it had just been handed, which only worked because no
union branch contains one. The query builder now always emits a WHERE, so
adding a branch with its own predicate cannot silently break the cursor.
2026-09-22 10:43:40 +01:00
arc53-machine 96a930c265 test: cover the new usage columns and the rollup exclusion
bucketed_totals now returns cost and cached_tokens, so the exact-equality
assertion in test_day_bucket_sums_per_day had to grow the two fields; it
pins that an unreported cache bin reads NULL rather than 0.

The per-user breakdown test was seeded with a source='schedule' row, which
is a run-level rollup and no longer counted. Reseeded on a real flow, and
both spend queries gained a test that a rollup is excluded.
2026-09-22 10:42:36 +01:00
arc53-machine c2d1893992 fix(admin): keyset export paging, rollup-free spend, consistent totals
Review pass over the preceding commits.

- The activity export paged by OFFSET. The journals are append-only and the
  feed is newest-first, so rows written mid-export shift the window down and
  repeat rows already emitted. Pages by keyset on (created_at, feed, id) now.
- top_token_users and the per-user breakdown counted run-level rollup rows.
  A scheduled run already has a row per LLM call, so its spend was billed
  twice -- invisible while the column was tokens, obvious once it was dollars.
  Both now exclude them, matching every other spend query.
- total_cost was summed from the per-model split, which drops rows with no
  model_id and so undercounted the figure printed above the series. Summed
  from the series instead.
- The CSV export encoded its detail cell without the fallback the NDJSON
  branch had, so a non-JSON-native value would have failed the stream.
- The activity view reset pagination in an effect, which fetched the stale
  page against the new filters before fetching again. Reset in the setters.
- The audit taxonomy is no longer re-exported from the helper module; the one
  caller that wanted it imports from where it lives.
2026-09-22 10:41:46 +01:00