* feat(cron): deterministic command payloads (run a shell command, no LLM)
Cron jobs always run an agent turn today, so deterministic work (health
probes, backups, syncs) pays model tokens on every fire. This adds a
"command" payload kind that runs a shell command directly in the gateway
process with zero model tokens, mirroring openclaw's command cron.
- store: CronPayload.Command (*CronCommandSpec — argv/cwd/env/input/
timeouts/output cap). Persists in the existing payload JSON blob, so
there is NO migration and no schema version bump.
- internal/cronexec: in-process runner with wall-clock + no-output
timeouts, per-stream output capping, and process-group termination so a
timed-out command's forked children are also killed.
- gateway_cron handler: command jobs run in-process and deliver stdout on
success (honoring the NO_REPLY sentinel). A non-zero exit / timeout
returns an error so the run is recorded as error and retried per
cron.max_retries; failures are NOT delivered, mirroring the agent path
(only successful output is announced — no channel spam).
- surfaces: cron.create RPC, the agent `cron` tool, and a new
`goclaw cron create` CLI all accept command payloads.
- security: gated by cron.command_enabled (default false). Commands run
with the gateway process's privileges, so the feature is opt-in per
gateway; when disabled the RPC and tool reject command payloads and the
handler refuses to run them.
- i18n (en/vi/zh), docs (08-scheduling-cron.md), and tests for the runner
and the handler command path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(cron): gate command payloads on the update surfaces too
handleUpdate (RPC + agent tool) passed CronJobPatch.Command straight to
UpdateJob, which switches the payload to command kind for any non-nil
Command — without the command_enabled gate or ValidateCronCommandSpec that
create enforces. A normal job could therefore be mutated into a command job
(or persisted with an invalid spec, e.g. empty argv) on a gateway where
command cron is disabled, breaking the disabled-gateway contract.
Both update surfaces now require cron.command_enabled and validate the spec
before UpdateJob, matching create. The agent tool parses the command via the
same path as add and drops the raw keys so a shell-string command can't break
the generic patch unmarshal. Regression tests added for RPC and tool update
(command disabled + invalid argv), plus a positive enabled-valid case.
Addresses review feedback from @mrgoonie on #1279.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
setupToolRegistry wires the filesystem tools' AllowPaths from the config/JSON5
default before ApplySystemConfigs overlays system_configs['allowed_paths'], so
DB-configured allowed paths never reached read_file / list_files / write_file /
edit / send_file. Only the rate limiter was re-applied after the overlay (#1111);
the AllowPaths analogue was missing, so agents were denied access to configured
shared directories outside their workspace even though the DB value was present
(visible as a read_file "access denied" log whose allowedPrefixes omit the
configured path).
Extract the user-allowed-path application into applyUserAllowedPaths and re-run it
from runGateway after the overlay, mirroring the rate-limiter re-apply. The helper
is idempotent (AllowPaths is additive and the prefix check is membership-based),
so the initial wiring call plus the re-apply is safe.
Test: cmd/gateway_tools_wiring_test.go asserts a path outside the workspace is
denied before the grant and readable after, that an unrelated path stays denied,
that repeated application is safe, and that an empty list is a no-op.
A run records its terminal run.status (completed/failed/cancelled) only
when the agent loop returns. If the gateway process is killed mid-run
(restart, crash, or SIGKILL after a slow graceful shutdown), the terminal
event is never emitted, so the run's last run.status item stays "started"
forever — it shows as perpetually running in the timeline and is never
counted as failed.
Mirror the cron scheduler's startup reset (recomputeStaleJobs flips stale
'running' jobs to 'interrupted'): add RunTimelineStore.RecoverInterruptedRuns,
invoked once on gateway startup. It finds runs that have a started run.status
but no terminal sibling and appends a terminal failed run.status item, marked
interrupted in metadata so it stays distinguishable from a genuine agent
failure. Running only at startup means nothing from before is still
executing, so there is no false-positive risk — unlike a periodic sweep
(see the deliberately-disabled tracing stale-trace loop).
- store: RecoverInterruptedRuns on the interface + PG and SQLite impls
- cmd/gateway_managed: invoke once after the timeline recorder is built
- tests: SQLite store test — orphan recovered, completed run untouched,
idempotent on re-run
The agent config UI (tools-profile-section) and i18n already expose a
per-agent "Rate Limit (per hour)" field saved into tools_config, but the
backend ignored it — the tool rate limiter only ever used the global
tools.rate_limit_per_hour. Wire the field through: ToolPolicySpec gains
RateLimitPerHour; the agent loop threads it into the tool-execution
context; the registry passes it to the limiter, which uses it in place
of the global max when > 0 (0 inherits the global).
The per-run session reset was gated on `!job.Stateless`, inverting the
flag: stateless jobs (the token-saving default) skipped the reset and
accumulated unbounded history, while stateful jobs were wiped every run.
Reset for `job.Stateless` instead, and clear BOTH session layers — the
goclaw session store AND the Claude CLI on-disk .jsonl. claude-cli
resumes its own session by a deterministic per-key UUID, so without
clearing the .jsonl a "stateless" run still replayed the entire
accumulated history (and silently grew it run after run).
Follow-up to #1260. A model (notably the claude-cli proxy under a
degraded session) sometimes emits a tool call as a bare
<invoke name="...">...</invoke> block with no <function_calls> wrapper.
fullToolCallBlockPattern only matched the wrapped form, so the bare
block slipped through to the tag-only strip, leaking the inner
<parameter> command text into the reply while the tool never ran.
Match and remove complete bare <invoke> blocks too (after the wrapped
form), add "<invoke name=" as a detection indicator, and log the
dropped tool name(s). Partial/unterminated artifacts still fall through
to the existing tag strip.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
stripGarbledToolXML removed only the tags of a model-emitted
<function_calls> block, leaking the inner <parameter> argument values
into the user-facing reply, and the dropped call was completely silent.
Detect a complete tool-call block, remove it whole, and log the
attempted tool name(s) at WARN so the no-op is diagnosable. Partial
artifacts still fall through to the existing tag-level strip.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Registering a remote MCP server (sse/streamable-http) whose hostname
resolves to a private IP is rejected at config-validation time by the
SSRF guard, with no production escape hatch -- the only bypass is a
test-only loopback flag. Self-hosted MCP servers on a private network
are therefore unregisterable, even though the runtime MCP client
connects to them fine (Test Connection succeeds and lists tools; only
create/update input validation blocks).
Add an opt-in, operator-configured allowlist (GOCLAW_MCP_ALLOWED_HOSTS,
empty by default) of trusted hostnames exempt from the private/loopback
IP block during MCP server URL validation only:
- security.ValidateAllowingHosts(url, allowedHosts): like Validate but
skips the private/loopback block for allowlisted hostnames. The
cloud-metadata/link-local (169.254/fe80), multicast and unspecified
ranges are never exempted, even for allowlisted hosts.
- mcp.SetAllowedHosts wires the operator allowlist into ValidateURL /
ValidateServerConfig; default empty => no behavior change.
- web_fetch / webhook / redirect SSRF paths are unchanged (those stay
agent-influenced and fully guarded).
Matching is case-insensitive on the pre-resolution hostname.
The BridgeToolNames allowlist omitted "delegate", so claude-cli /
Claude-Code-backed agents never received the tool even when their
orchestration mode resolved to "delegate" with active agent_links.
The native agent loop injects delegate per orchestration mode, but
bridge sessions only see this curated allowlist, making inter-agent
delegation impossible from bridge-backed agents.
delegate is safe to expose unconditionally: it self-gates via
CanDelegate/agent_links and resolves its source agent from the
X-Agent-ID header context, exactly like team_tasks (already bridged).
The sessions.reset RPC only reset the native session store, leaving the
Claude CLI-backed history (.jsonl + CLAUDE.md) in place. The /reset chat
command already clears it via providers.ResetCLISession, so RPC callers
got an incomplete reset for claude-cli agents — stale history was resumed
on the next turn.
Mirror the chat path by calling ResetCLISession from handleReset. It is a
no-op when the CLI provider is unused. The call is wired through an
overridable package var so the behavior can be regression-tested.
makeCronJobHandler builds the agent RunRequest with no provider/model, so cron
jobs always run on the agent default provider. High-frequency scheduled jobs
(e.g. triage) can't be routed to a cheaper model the way heartbeats already can
via agent_heartbeats.provider_id/model.
Add provider_id/model to cron_jobs (Postgres migration 000074; SQLite schema +
migration v42→43, both mirroring agent_heartbeats) and resolve them into
RunRequest.ProviderOverride/ModelOverride in the handler, reusing the existing
per-run override plumbing that heartbeats use. NULL/unset → agent default, so
existing jobs are unaffected.
Settable today via UpdateJob patch (CronJobPatch.ProviderID/Model) or directly
in the DB; tool/RPC/dashboard exposure left as a follow-up.
Verified: go build ./... and go build -tags sqlite ./... (exit 0), go vet,
gofmt, and cron unit tests on both pg and sqlite stores.
bridgeContextMiddleware injects the agent UUID (store.WithAgentID) but never the
agent key, so session tools (sessions_list/history/send/status) that resolve the
caller via tools.ToolAgentKeyFromCtx always get "" and fail with "agent context
required" when invoked over /mcp/bridge (e.g. by a claude-cli provider agent).
Session keys are namespaced by agent key (agent:<key>:...), not UUID, and the
bridge has no RunContext fallback.
Inject the key in the same block that already fetches the agent for shell-deny
overrides — symmetric with store.WithShellDenyGroups, no extra DB call, no new
imports.
#1094 makes the same fix at this site but bundles it into a large, stale
(CONFLICTING) ACP/i18n refactor; this is the minimal extraction of that slice.
It omits #1094's companion store.WithAgentKey write, which nothing on dev reads
(every session tool reads tools.ToolAgentKeyFromCtx).
Add internal/gateway/bridge_context_test.go: a regression guard asserting a
signed X-Agent-ID lands the agent key in tools.ToolAgentKeyFromCtx, plus a
no-store negative control.
Signed-off-by: Jaegeon Oh <zezaeoh@gmail.com>