Runs on models without a registered tokenizer (e.g. 9router brand models)
ended with the generic "Agent couldn't generate a response" fallback even
though the real request used about 55% of the context window.
PruneStage counted history with TokenCounter, which falls back to a
chars/2 heuristic for unregistered models and overcounted about 1.8x.
Once over budget it ran memory flush (~35s, invisible in traces), then
mid-loop compaction, which cannot summarize a history made only of tool
call/result pairs. The callback reported the untouched history as
compacted, PruneStage still saw it over budget and returned AbortRun
before any LLM call, and FinalizeStage replaced the empty reply with the
fallback.
- PruneStage and ContextStage overhead count with the request guard's
BudgetCounter. PruneStage no longer controls loop flow; the final
request guard in ThinkStage decides.
- CompactMessages returns ErrNotCompacted when history is unchanged.
Callers stop counting it as a compaction and do not retry it in the
same run, while post-run summarization still sees the pressure.
- When the guard exhausts every reduction step, ThinkStage stops the run
with a localized chat.context_budget_exceeded notice instead of an
error, so the run's tool results are still persisted. The stop reason
marks the trace and agent span as error; team tasks, cron and
heartbeat treat it as a failure via RunOutcome.Failure().
- Memory flush and mid-loop compaction emit event spans.
- Web and desktop UIs treat an unset context_pruning as enabled (the
backend default since 7639a8c0), keep it unset when untouched, and can
re-enable pruning after it was turned off.
vault_search only indexed title + path + the auto-summary, and the summary
is written from the first 3000 runes of the file. Anything the summary left
out (room names, prices, codes) could never be found. On top of that the
FTS query used plainto_tsquery, which ANDs every word, so a normal question
like "what is the price of the Deluxe Ocean Suite?" matched nothing even when
the key words were indexed.
- Add vault_document_chunks (migration 98): the file body split into
chunks with their own tsvector and embedding. body_indexed_hash on
vault_documents records which content_hash the chunks came from, so
unchanged files are not re-chunked or re-embedded.
- The enrich worker rebuilds chunks when a file changes. This needs no LLM,
so it runs even when no provider is configured.
- Rescan backfills chunks for docs indexed before this change, since they
already have a summary and never go back through the worker.
- FTS matches any query word; ts_rank still ranks docs with more matching
words first. Both FTS and vector search look at the doc and its chunks
and score each doc by its best hit.
SQLite is unchanged: its vault search is LIKE on title/path only.
Register Requesty (https://router.requesty.ai/v1) next to OpenRouter on the
existing OpenAI-compatible transport: config and env vars
(GOCLAW_REQUESTY_API_KEY, GOCLAW_REQUESTY_BASE_URL), secret masking, DB and
in-memory registration, CLI setup, doctor, placeholder provider, OpenAPI enum,
Web/Desktop provider lists and docs.
The models list merges Requesty managed policies (GET /models/managed) with
the key's catalog from GET /models.
Map skills/anysearch to the AnySearch integration review checklist
(basic built-in + extensions). Record 2026-09-24 anonymous live smoke
for search, get_sub_domains, batch_search, extract, and vertical search.
Drop the Sh/PS ports in response to PR triage feedback on the
maintenance/security surface. Keep Python (requests) as primary and
Node as the zero third-party-dep fallback. Docs and offline tests
updated accordingly.
Add skills/anysearch (vendored from anysearch-ai/anysearch-skill v3.1.1)
so agents get general/anonymous search, parallel batch search, vertical
domain search, and full-page extract. Includes MAINTENANCE.md and
docs/anysearch-skill.md. No internal/, ui/, or migration changes.
Disclosure: AnySearch Open Source Bounty Claim goclaw#001.
Creating a predefined agent with a description kicks off summoning in a
background goroutine. It finishes 10-20 seconds after the create response and
writes SOUL.md, IDENTITY.md, CAPABILITIES.md and the agent frontmatter.
For a client that manages agents as code that is data loss. Such a client
brings its own context files and writes them through agents.files.set as soon
as create returns; summoning lands afterwards and replaces them. Nothing
fails: no error from the API, nothing in the logs, and the agent quietly runs
on generated text instead of the committed role. The only way to avoid it
today is to poll the agent status until it leaves "summoning" — which means
every automation has to know that summoning exists at all.
POST /v1/agents now accepts an optional "summon" field. Absent or true keeps
today's behaviour, so existing clients are unaffected. False creates the agent
active with the seeded template files and starts no LLM call; summoning is
still available afterwards via POST /v1/agents/{id}/resummon.
The status fallback also now resets an explicitly requested "summoning" status
when nothing will summon the agent: that status is only ever cleared by the
summoner, so such an agent would otherwise stay stuck in it forever.
Surface parity: CLI and WS RPC (agents.create) create agents without a
description and never summoned, so they need no flag; the web UI create flow
wants summoning and keeps the default.
- docs: device.pair.update and the `permanent` option on approve in
docs/04-gateway-protocol.md, docs/19-websocket-rpc.md and
websocket-protocol.md; the paired-device TTL row in docs/09-security.md
now mentions the admin opt-out.
- store.ErrPairedDeviceNotFound: SetPairingPermanent wraps it in both stores.
device.pair.update maps it to NOT_FOUND and any other store error to
INTERNAL, so a DB failure no longer reads as "not found".
- web UI: approve and make-permanent/set-expiry now toast the server error
and reload the list in `finally`. A partially applied approve (paired, but
the permanent write failed) shows up in the table instead of leaving the
dialog dead-ended.
- SQLite ListPaired: a stored expiry that fails to parse stays 0 (expires,
date unknown) rather than being mistaken for permanent; the UI renders it
as "--" instead of a 1970 date.
Tests: gateway handler error mapping (NOT_FOUND / INTERNAL / OK), sentinel
checks in the PG and SQLite store tests, SQLite unreadable-expiry case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
read_video called with a url parameter always failed under an agent budget:
tool:read_video: cannot verify streamed native media against the agent
context budget (no in-memory payload to count); refusing to send
read_audio and read_document appeared to work but only for very small files.
Both symptoms share one root cause.
5ca433b8 introduced the complete-input invariant and, to keep out-of-band
media honest, appended the standard-base64 encoding of the payload as a
synthetic guard-only message. That message is counted as text. Measured against
the bundled BudgetCounter, base64 costs 0.956 tokens per raw byte, so an 11.6 MB
video counted 11,161,058 tokens and exceeded a 200k, a 1M and a 2M window alike.
The practical ceiling was roughly 200 KB at a 200k window while videoMaxBytes is
100 MB. The read_video URL transport streams to the File API without buffering,
so it had no bytes to count at all and was routed to a helper that refused to
send whenever an agent budget was present.
Media is now priced by what a provider actually bills for it. Every number below
is either a published per-unit rate or a published provider limit; none is
derived from a byte size, because a static-image video compresses arbitrarily
small and no bitrate floor exists. An earlier revision of this branch tried one
and a 120-second, 20,627-byte clip priced at 526 tokens against a true 31,560.
internal/mediabudget prices one payload:
- Video and audio: ffprobe measures duration. 263 tokens/second for video,
32 for audio, both published for static processing at 1 FPS.
- PDF: pdfinfo counts pages at 258 tokens/page. When the page count cannot be
read, the charge is the proven ceiling of 1000 pages, which is the most a
provider will accept and therefore the most it can bill.
- Video and audio that cannot be measured are refused. Upstream already
refused unverifiable native media for every URL on the Gemini streamed path;
this narrows that refusal from every URL to only what genuinely cannot be
measured, rather than removing it. PDF differs because its page count is
cheaply measurable and its ceiling is small enough to stay usable.
A remote video is measured without downloading it: two ranged GETs, 512 KB from
the head and 512 KB from the tail, written into a sparse temp file sized to the
declared total and handed to ffprobe. The tail matters because every container
that puts its index at the end keeps it there: head-only probing under-reports
mpeg by 98% and ogg by 79%, and a plain prefix makes ffprobe under-report a
30-second WAV as 0.74 seconds because it clamps to the bytes it can see. Sizing
the temp file to the real total fixes that. Verified end to end on an 11,673,105
byte MP4 served by nginx: 1 MB of ranged reads yielded duration 40.000000,
identical to ffprobe reading the whole URL, for a charge of 10,520 tokens. Those
requests reuse the existing SSRF-safe path, security.WithPinnedIP plus
security.NewSafeClient(0), and no URL is ever handed to an external binary.
Beyond the reported bug, two pre-existing gaps let large media reach a provider
almost unpriced. ExecuteWithChain treated every callProvider error as a provider
failure and advanced to the next entry, and the non-Gemini branches of read_video
and read_document reserved without pricing their payload at all. Measured under a
20,000-token window before this change, a 40 MB video and a 1000-page PDF each
reached a provider charged about 1,600 tokens. Budget refusals are now terminal
in the chain and every media branch prices its payload, so both reach no provider
at all. Genuine provider failures still fail over.
Known limits, stated rather than discovered:
- The /Type /Page scan that guards against a forged /Count is a floor, not a
bound. Pages inside a compressed object stream are invisible to it, and a
9,484-byte PDF built that way is charged 258 tokens for 1000 pages. A real
pdfinfo reads such files correctly; the scan only ever raises a probed count.
- A video URL whose origin does not serve byte ranges is now refused on the
non-Gemini path too, and the refusal is terminal. Upstream forwarded such URLs
unpriced. A HEAD giving only a size is not enough to price one.
- read_video and read_audio require ffprobe. Docker images install it by default
except the base variant; bare binaries and the desktop build do not ship it.
- A hostile origin can craft a container ffprobe reads as about one second.
Reservation.Reconcile overwrites the estimate with the provider's reported
usage, so this weakens the gate rather than defeating it.
Byte ceilings videoMaxBytes, audioMaxBytes and documentMaxBytes are unchanged.
No new module dependency; ffprobe and pdfinfo are optional runtime probes.
* feat(teams): let a human cancel and retry stuck team tasks from the dashboard
A task that ended up blocked, stale, failed or cancelled could not be
recovered from the UI: the dashboard only deletes terminal tasks and
approves/rejects in_review ones. CancelTask and ResetTaskStatus existed
in the store but were reachable only through the lead agent's team_tasks
tool, which needs the full task UUID the dashboard never shows (#506).
Two WS RPCs, wired to Retry / Cancel buttons in the task detail dialog:
- teams.tasks.cancel — any task not yet completed/cancelled. Optional
reason is stored as the result and posted as a comment so the lead
sees it on the board. Dependents are unblocked by the store; they are
not dispatched here because the dashboard has no agent turn.
- teams.tasks.retry — stale / failed / cancelled / in_review, plus
blocked tasks that nothing blocks any more. A human comment is
required: it is posted on the task and appended to the assignment
prompt, so the assignee gets the missing answer in the same message.
Optional agentId reassigns; the lead is refused as assignee (same
guard as teams.tasks.assign).
ResetTaskStatus (pg + sqlite) now also accepts blocked, guarded on the
handler side by an empty blocked_by. Handler tests cover both actions
with a stub store: reason/comment persistence, required comment,
status gating, blocked_by guard, lead guard, cross-team IDOR.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(permissions): classify teams.tasks.cancel and teams.tasks.retry as write methods
Every WS method must be classified in internal/permissions/policy.go,
otherwise the router answers UNAUTHORIZED for every role. Caught on a
live gateway (the drift test TestMethodRole_DriftCoverage… flags it, I
had not run that package). Both sit next to approve/reject/assign as
operator-level write methods; the write-methods test now pins them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat(teams): let the retry dialog pick the assignee
A cancelled or stale task may have no owner (created unassigned, or the
member was removed). teams.tasks.retry then needs agentId, which the
dialog did not send, so Retry would fail with "task has no assignee".
The retry dialog now shows an assignee select (team members minus the
lead, defaulting to the current owner) and passes agentId only when it
differs from the owner. Retry is offered only when there is someone to
pick.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
channel_pending_messages is a buffer, not an archive. Rows were deleted
outright on two paths: the bot being mentioned hands the buffer to the
agent and clears the key, and LLM compaction replaces old rows with a
summary. For group capture that buffer held the only copy of the raw
text, so a single mention or compaction pass destroyed days of messages
nothing had read yet.
Copy every row into channel_message_archive inside the same transaction
as the delete, tagged with the reason (consumed, compacted, stale).
Archived rows keep their original id, so a replayed delete is a no-op.
Add ListArchivedByKey so consumers that need full history read the
archive instead of the buffer.
A single dispatch goroutine delivered every reply in the process — every
channel, every user — each Send blocking the next. One media upload of a
few seconds froze all outbound traffic for that long, and sustained
throughput was capped at one Send round-trip regardless of lane capacity.
The ordering that loop actually protected is per-conversation: a run emits
block replies, retry notices and a final answer to one ChatID, and those
must arrive in that order. A global order was never required.
Dispatch now shards on channel+ChatID. A conversation always maps to the
same shard and is delivered serially there, so the ordering guarantee is
unchanged, while unrelated conversations proceed in parallel.
GOCLAW_OUTBOUND_SHARDS tunes the worker count (default 8).
Temp media needed a real fix to go with it. Serial dispatch deduped it
implicitly — the first delivery removed the file, and a later message
carrying the same path found os.Stat failing and skipped it. Across
concurrent shards that check alone is a TOCTOU race, so claims are now
taken atomically and released after the send.
Also makes the pool limits that bound this path configurable, defaults
unchanged: GOCLAW_PG_MAX_OPEN_CONNS / GOCLAW_PG_MAX_IDLE_CONNS, and
GOCLAW_HTTP_MAX_IDLE_CONNS / GOCLAW_HTTP_MAX_IDLE_CONNS_PER_HOST. The
per-host limit binds when one provider takes nearly all traffic: every
request past it pays a fresh TCP+TLS handshake, which a deployment
raising LaneMain needs to be able to lift.
Show an ephemeral "agent is working" indicator (thinking/searching/
generating/analyzing…) in Bitrix24 chat while the agent processes, so
users on this non-streaming channel aren't left staring at silence
until the final reply. Send behavior is unchanged; no LLM call, no DB.
Generic layer (internal/channels):
- New optional ActivityIndicatorChannel interface.
- HandleAgentEvent routes run.started->THINKING, tool.call->mapped
status, tool.result->ANALYZING; static tool->status mapping.
- Conditional heartbeat ticker fills LLM-inference gaps (re-sends only
when idle), 5s per-run throttle caps call rate; ticker stopped on
terminal events AND in UnregisterRun (safety net for missed terminals).
Bitrix24 (internal/channels/bitrix24):
- OnActivityEvent calls imbot.v2.Chat.InputAction.notify best-effort,
drop-on-limit (raw Client.Call, no retry) so cosmetic notifies never
steal leaky-bucket capacity from real message sends.
- Per-channel activity_indicator toggle (default on).
Web UI + docs: dashboard toggle, i18n en/vi/zh, channels doc section.
Tests: 26 unit tests incl. UnregisterRun-stops-ticker regression.
* docs(i18n): design spec for Russian (ru) language support
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(i18n): register ru locale constant and IsSupported
* feat(i18n): scaffold ru catalog (en placeholders)
* feat(i18n): translate ru backend catalog to Russian
* fix(i18n): polish ru catalog grammar per review
* feat(i18n): wire ru into system messages, config enum, parity test
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(i18n): add ru to web language constants and system-message locales
* feat(i18n): scaffold ru web locale dir (en placeholders)
* feat(i18n): wire ru resources into web i18next config
* feat(i18n): translate ru web locale to Russian (41 namespaces)
* fix(i18n): normalize ru web terminology (арендатор, привязка)
* fix(i18n): restore dropped 'named' in ru setup OAuth hint
* feat(i18n): wire ru into desktop i18next and language pickers
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(i18n): add ru desktop locale (Russian, 16 namespaces)
* test(i18n): extend traces/hooks locale guards to cover ru
* chore(deploy): add repeatable fork-dev local build+deploy tooling
Adds a one-command, idempotent way to run goclaw from a locally built
image that merges fresh upstream dev with un-merged fork features
(Telegram trigger-words PR #1383 + Russian locale), independent of the
published :latest image.
- scripts/deploy-forkdev.sh: fetch origin/dev -> merge feature branches
into local/fork-dev worktree -> push fork/dev -> build -> up -d.
- docker-compose.forkdev.yml: overlay pinning image to goclaw:forkdev
with pull_policy:never so :latest never silently replaces the build.
- .gitignore: track the deploy helper (past deploy-*.sh), ignore the
../.goclaw-forkdev integration worktree.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(telegram): agent-callable message reaction (message action=react)
Adds a "react" action to the message tool so an agent can set an emoji
reaction on an existing message (e.g. mark its own status post 👍 once
fully paid). Mirrors the ChannelEditor wiring: ReactionSetter capability
-> Manager.ReactToMessage -> telegram Channel.ReactToMessage (validates
against Telegram's supported reaction set; ✅/❌ are rejected).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(config): remove ignored session scoping controls (#1405)
Co-authored-by: Collective Developer <man@collective.dev>
* feat(telegram): agent-callable message reaction (message action=react) (#1407)
Adds a "react" action to the message tool so an agent can set an emoji
reaction on an existing message (e.g. mark its own status post 👍 once
fully paid). Mirrors the ChannelEditor wiring: ReactionSetter capability
-> Manager.ReactToMessage -> telegram Channel.ReactToMessage (validates
against Telegram's supported reaction set; ✅/❌ are rejected).
Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* ci: build & publish fork image to GHCR (temporary, pre-upstream)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(memory): use 'russian' FTS config for memory_chunks recall
Memory keyword search (memory_search.go FTS + the generated tsv column)
used the 'simple' text-search config, which does no stemming. Russian
grammatical cases therefore never matched: a genitive query token
("Дианы") could not hit the stored nominative token ("Диана"), so
memory_search returned nothing for natural queries even though the fact
was saved. This is the sole recall path on deployments without an
embedding provider (ChatGPT-OAuth/codex cannot embed), so recall was
effectively dead there.
Switch both sides to the Snowball 'russian' config so index and query
lexemes stem to the same root ("диан"), restoring case-insensitive
recall. Migration 000094 recreates the generated tsv column (a GENERATED
expression cannot be ALTERed in place; ADD COLUMN backfills all rows).
Surface parity: desktop/SQLite N/A — SQLite memory search uses LIKE, a
separate mechanism unaffected by the PG regconfig.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: ナム <namnhfreelancer@gmail.com>
Co-authored-by: Collective Developer <man@collective.dev>
- Fix image_generation nil-pointer crash on codex agents: the sentinel now
carries a name-only Function so the many sites reading td.Function.Name never
nil-deref; codex_build still branches on Type.
- Agent-declared trigger words in IDENTITY.md wake the bot in groups without an
@mention (whole-word, Cyrillic-aware; text + caption), cached per-agent 60s.
- channel_post support with a synthetic sender + a recover() guard so a
malformed update can't crash the gateway.
- message tool action=edit (editMessageText + editMessageCaption fallback),
targeting the replied-to message via reply_to_message_id.
- message tool topic=<name> posts into a named forum topic; topics learned from
forum_topic_created into channel_contacts and resolved by name.
Co-authored-by: skensel <skensel@MacBook-Pro-skensel.local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Surface cache-read and cache-creation input tokens (plus the
prompt_tokens_include_cached_segments flag) in webhook usage responses,
so callers can distinguish cached from non-cached input tokens.
The data already flowed end-to-end via providers.Usage and
agent.RunResult.Usage; the two webhook envelope structs (webhookLLMUsage
for the sync LLM webhook, callbackUsage for the async callback payload)
copied only 3 of the fields, dropping the cache data. Add the 3 cache
fields (mirroring providers.Usage JSON tags exactly, all omitempty) and
populate them. The async callbackUsage mapping is extracted into a pure
newCallbackUsage helper for unit testing.
Additive and non-breaking: omitempty keeps responses without caching
byte-identical. No schema change.
fix(usage): repair cost analytics and display precision
- Automatic OpenRouter pricing sync with cost backfill for traces/snapshots/events
- Live usage data merging for current-hour dashboard accuracy
- 2-decimal API cost formatting across usage/overview pages
- Comprehensive test coverage (PG, SQLite, HTTP, UI)
Merged by github-maintain automation.
Lightpanda upstream merged the AX-tree nodeId fix in
lightpanda-io/browser#2232. The TestLightpanda_KnownUpstreamGaps canary
fired on the latest image, so:
- Rename to TestLightpanda_Snapshot_AndEval and assert the snapshot
returns refs + non-empty text (instead of asserting it fails).
- Flip AX-snapshot in the compatibility matrix from gap to fully
supported. Replace the "Known upstream gap" section with a
"Minimum Lightpanda version" note pointing at the upstream PR.
- Eval works on Lightpanda when called with go-rod's expected function
form (e.g. "() => document.title"); only bare expressions fail, and
they fail on Chrome too. Fix the integration gap test and docs that
incorrectly attributed this to a Lightpanda bug.
- docker-compose.lightpanda.yml: add the missing
command: lightpanda serve ... — the image's entrypoint isn't
lightpanda, so without an explicit command the sidecar wouldn't
start.
Live testing surfaced three issues:
- Lightpanda numbers targets per-browser, and each conn is its own
browser, so every tab gets the same upstream targetID
("FID-0000000001"). Synthesize globally-unique "lp-N" keys for our
internal map so multi-tenant tab tracking works.
- page.Info() returns valid data once post-open then errors on
subsequent calls, which made ListTabs silently drop tabs. Cache URL
and Title at OpenTab time and read from the cache in ListTabs.
- rod.Browser.Close() calls Browser.close which Lightpanda doesn't
implement; the WS drops cleanly anyway. Swallow the error to quiet
the noisy log line.
Two Lightpanda upstream bugs are documented in docs/browser-backends.md
and exercised by TestLightpanda_KnownUpstreamGaps:
1. Accessibility.getFullAXTree returns nodeId as a JSON number
(CDP spec: string)
2. Runtime.evaluate rejects go-rod's function-apply wrapper
- docker-compose.lightpanda.yml: opt-in overlay running
lightpanda/browser:latest, wires GOCLAW_BROWSER_REMOTE_URL and
GOCLAW_BROWSER_BACKEND so the manager picks the right code path.
- docs/browser-backends.md: compatibility matrix vs Chrome (screenshot,
multi-tab, cookie sharing, etc.) and guidance on when to pick which.
* fix(feishu): require exact bot_open_id match for mention detection
- Previously, when bot_open_id was empty, ALL mentions were treated as bot mentions
- This caused multiple agents in same group to all respond to any mention
- Now requires bot_open_id to be set AND match exactly for mentionedBot=true
- Fixes issue where CPPAI PM would respond even when other agent was mentioned
* fix(feishu): gate group pairing by target bot
---------
Co-authored-by: ntduc <ntduc@cpp.ai.vn>
* feat(cron): deterministic command payloads (run a shell command, no LLM)
Cron jobs always run an agent turn today, so deterministic work (health
probes, backups, syncs) pays model tokens on every fire. This adds a
"command" payload kind that runs a shell command directly in the gateway
process with zero model tokens, mirroring openclaw's command cron.
- store: CronPayload.Command (*CronCommandSpec — argv/cwd/env/input/
timeouts/output cap). Persists in the existing payload JSON blob, so
there is NO migration and no schema version bump.
- internal/cronexec: in-process runner with wall-clock + no-output
timeouts, per-stream output capping, and process-group termination so a
timed-out command's forked children are also killed.
- gateway_cron handler: command jobs run in-process and deliver stdout on
success (honoring the NO_REPLY sentinel). A non-zero exit / timeout
returns an error so the run is recorded as error and retried per
cron.max_retries; failures are NOT delivered, mirroring the agent path
(only successful output is announced — no channel spam).
- surfaces: cron.create RPC, the agent `cron` tool, and a new
`goclaw cron create` CLI all accept command payloads.
- security: gated by cron.command_enabled (default false). Commands run
with the gateway process's privileges, so the feature is opt-in per
gateway; when disabled the RPC and tool reject command payloads and the
handler refuses to run them.
- i18n (en/vi/zh), docs (08-scheduling-cron.md), and tests for the runner
and the handler command path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cron): gate command payloads on the update surfaces too
handleUpdate (RPC + agent tool) passed CronJobPatch.Command straight to
UpdateJob, which switches the payload to command kind for any non-nil
Command — without the command_enabled gate or ValidateCronCommandSpec that
create enforces. A normal job could therefore be mutated into a command job
(or persisted with an invalid spec, e.g. empty argv) on a gateway where
command cron is disabled, breaking the disabled-gateway contract.
Both update surfaces now require cron.command_enabled and validate the spec
before UpdateJob, matching create. The agent tool parses the command via the
same path as add and drops the raw keys so a shell-string command can't break
the generic patch unmarshal. Regression tests added for RPC and tool update
(command disabled + invalid argv), plus a positive enabled-valid case.
Addresses review feedback from @mrgoonie on #1279.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Server-side webhook agent runs (sync, async, and admin test) used Stream:false,
so OpenAI-compatible routers that only cache streaming requests never populated
or served their prompt cache. Webhook runs paid full input-token price on every
turn even with a stable session and an identical multi-turn prefix, while WS chat
(Stream:true) got cache hits from the 3rd turn.
Add gateway.webhook_stream (default true, env GOCLAW_WEBHOOK_STREAM) and apply it
to all three run sites. ChatStream returns the fully assembled response, so the
payload returned to the caller is unchanged. Set to false to restore non-streaming.
Add total-count pagination to the webhook admin endpoints and the web UI.
Store:
- WebhookStore/WebhookCallStore gain Count; WebhookListFilter gains
IncludeRevoked + Query (PG + SQLite, parameterized, tenant-scoped)
API:
- GET /v1/webhooks and GET /v1/webhooks/{id}/calls return
{items, total, limit, offset} with server-side search + revoked filtering
Web UI:
- server-driven list pager + search/revoked filter; call-history dialog uses
the real total (fixes the full-page "has more" boundary bug)
- i18n pager labels (en/vi/zh)
Tests: store pagination integration test; mock stores implement Count.
No schema migration (read-only COUNT).