6 Commits
Author SHA1 Message Date
thotam ce9e7af3fb fix(tools): price out-of-band media by measured duration, not by its bytes (#1564)
read_video called with a url parameter always failed under an agent budget:

    tool:read_video: cannot verify streamed native media against the agent
    context budget (no in-memory payload to count); refusing to send

read_audio and read_document appeared to work but only for very small files.
Both symptoms share one root cause.

5ca433b8 introduced the complete-input invariant and, to keep out-of-band
media honest, appended the standard-base64 encoding of the payload as a
synthetic guard-only message. That message is counted as text. Measured against
the bundled BudgetCounter, base64 costs 0.956 tokens per raw byte, so an 11.6 MB
video counted 11,161,058 tokens and exceeded a 200k, a 1M and a 2M window alike.
The practical ceiling was roughly 200 KB at a 200k window while videoMaxBytes is
100 MB. The read_video URL transport streams to the File API without buffering,
so it had no bytes to count at all and was routed to a helper that refused to
send whenever an agent budget was present.

Media is now priced by what a provider actually bills for it. Every number below
is either a published per-unit rate or a published provider limit; none is
derived from a byte size, because a static-image video compresses arbitrarily
small and no bitrate floor exists. An earlier revision of this branch tried one
and a 120-second, 20,627-byte clip priced at 526 tokens against a true 31,560.

  internal/mediabudget prices one payload:
  - Video and audio: ffprobe measures duration. 263 tokens/second for video,
    32 for audio, both published for static processing at 1 FPS.
  - PDF: pdfinfo counts pages at 258 tokens/page. When the page count cannot be
    read, the charge is the proven ceiling of 1000 pages, which is the most a
    provider will accept and therefore the most it can bill.
  - Video and audio that cannot be measured are refused. Upstream already
    refused unverifiable native media for every URL on the Gemini streamed path;
    this narrows that refusal from every URL to only what genuinely cannot be
    measured, rather than removing it. PDF differs because its page count is
    cheaply measurable and its ceiling is small enough to stay usable.

A remote video is measured without downloading it: two ranged GETs, 512 KB from
the head and 512 KB from the tail, written into a sparse temp file sized to the
declared total and handed to ffprobe. The tail matters because every container
that puts its index at the end keeps it there: head-only probing under-reports
mpeg by 98% and ogg by 79%, and a plain prefix makes ffprobe under-report a
30-second WAV as 0.74 seconds because it clamps to the bytes it can see. Sizing
the temp file to the real total fixes that. Verified end to end on an 11,673,105
byte MP4 served by nginx: 1 MB of ranged reads yielded duration 40.000000,
identical to ffprobe reading the whole URL, for a charge of 10,520 tokens. Those
requests reuse the existing SSRF-safe path, security.WithPinnedIP plus
security.NewSafeClient(0), and no URL is ever handed to an external binary.

Beyond the reported bug, two pre-existing gaps let large media reach a provider
almost unpriced. ExecuteWithChain treated every callProvider error as a provider
failure and advanced to the next entry, and the non-Gemini branches of read_video
and read_document reserved without pricing their payload at all. Measured under a
20,000-token window before this change, a 40 MB video and a 1000-page PDF each
reached a provider charged about 1,600 tokens. Budget refusals are now terminal
in the chain and every media branch prices its payload, so both reach no provider
at all. Genuine provider failures still fail over.

Known limits, stated rather than discovered:
- The /Type /Page scan that guards against a forged /Count is a floor, not a
  bound. Pages inside a compressed object stream are invisible to it, and a
  9,484-byte PDF built that way is charged 258 tokens for 1000 pages. A real
  pdfinfo reads such files correctly; the scan only ever raises a probed count.
- A video URL whose origin does not serve byte ranges is now refused on the
  non-Gemini path too, and the refusal is terminal. Upstream forwarded such URLs
  unpriced. A HEAD giving only a size is not enough to price one.
- read_video and read_audio require ffprobe. Docker images install it by default
  except the base variant; bare binaries and the desktop build do not ship it.
- A hostile origin can craft a container ffprobe reads as about one second.
  Reservation.Reconcile overwrites the estimate with the provider's reported
  usage, so this weakens the gate rather than defeating it.

Byte ceilings videoMaxBytes, audioMaxBytes and documentMaxBytes are unchanged.
No new module dependency; ffprobe and pdfinfo are optional runtime probes.
2026-09-12 10:31:12 +07:00
Clark Cant 5ca433b8ac fix: persist mid-loop compaction to stop re-compaction loop and revive episodic
The v3 pipeline compacts session history mid-loop (prune_stage +
final-request guard) but only mutates the run's message buffer, never the
session store. Each turn reloads full history and re-compacts from scratch:
message_tokens climb 129k->156k across turns while every turn compacts back
down to ~60k. The lossy compaction differs per run, degrading the agent.

The same missing persistence stalls episodic memory: the cumulative
compaction count never advances, so the episodic worker's idempotency key
(sessionKey:count) is pinned and every cycle after the first is skipped.
Observed on live traffic: 8 run.completed since deploy, 0 new episodic.

Fixes, all reusing existing machinery (no new store methods, no migrations):

- Bug A: emitSessionCompleted reads cumulative GetCompactionCount (matching
  the legacy v2 path) instead of the per-run counter that resets to 0.
- Bug B/anti-loop: finalize passes state.Prune.MidLoopCompacted into
  maybeSummarize; under pressure it lowers the trigger to a unit-aligned
  threshold (compactionInputCap - overhead, same MaxRequestShare the guard
  uses) so the compaction is PERSISTED via the existing TruncateHistory +
  IncrementCompaction path. Defensive floor prevents over-compaction on
  pathological config; tool-result-only bloat still skips (history-only).
- Bug C: SourceID embeds the count (sessionKey:count) so the eventbus dedup
  key advances per compaction cycle instead of swallowing rapid same-session
  turns within the 5m TTL.

Tests: episodic compaction, maybe_summarize pressure, request budget.
go build (PG + sqliteonly), go vet, go test -race all green.
2026-07-25 00:39:02 +07:00
Duc Nguyen 747b58d24c fix(usage): repair cost analytics and display precision (#1330)
fix(usage): repair cost analytics and display precision

- Automatic OpenRouter pricing sync with cost backfill for traces/snapshots/events
- Live usage data merging for current-hour dashboard accuracy
- 2-decimal API cost formatting across usage/overview pages
- Comprehensive test coverage (PG, SQLite, HTTP, UI)

Merged by github-maintain automation.
2026-07-03 06:24:21 +07:00
Goon d64a31ebdb fix(usage): enforce caps on auxiliary llm calls 2026-05-24 11:13:03 +07:00
Goon e4acb147d6 feat(usage): edit cap policies and trace decisions 2026-05-23 09:34:29 +07:00
Goon 9ff5a629d2 feat(usage): add AI budget usage caps 2026-05-23 00:16:03 +07:00